arxiv.org/abs/2501.04040v2
The rapid advancement of artificial intelligence, particularly with the development of Large Language Models (LLMs) built on the transformer architecture, has redefined the capabilities of natural language processing. These models now exhibit remarka...
arxiv.org/abs/2306.08302v3
Large language models (LLMs), such as ChatGPT and GPT4, are making new waves in the field of natural language processing and artificial intelligence, due to their emergent ability and generalizability. However, LLMs are black-box models, which often...
arxiv.org/abs/2404.05904v2
Large Language Models (LLMs) have transformed the Natural Language Processing (NLP) landscape with their remarkable ability to understand and generate human-like text. However, these models are prone to ``hallucinations'' -- outputs that do not align...
arxiv.org/abs/2303.10431v1
Large pre-trained vision-language models (VLMs) reduce the time for developing predictive models for various vision-grounded language downstream tasks by providing rich, adaptable image and text representations. However, these models suffer from soci...
arxiv.org/abs/2510.19318v1
The increasing reliance on natural language generation (NLG) models, particularly large language models, has raised concerns about the reliability and accuracy of their outputs. A key challenge is hallucination, where models produce plausible but inc...
arxiv.org/abs/2206.04615v3
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future resea...
github.com/Siddhhika/Lok-Sabha-Elections-Data-Analysis
The recently concluded 2024 Lok Sabha elections provided a wealth of data, offering deep insights into the current political landscape of India. The data was derived from the detailed election data available on the Election Commission of India's results page.…
arxiv.org/abs/2410.15547v1
Data cleaning is a crucial yet challenging task in data analysis, often requiring significant manual effort. To automate data cleaning, previous systems have relied on statistical rules derived from erroneous data, resulting in low accuracy and recal...
arxiv.org/abs/2101.07523v1
For some scientific questions, empirical data are essential to develop reliable simulation models. These data usually come from different sources with diverse and heterogeneous formats. The design of complex data-driven models is often shaped by the...
github.com/Stu-Vic/sql-challenge
It is a beautiful spring day, and it is two weeks since you have been hired as a new data engineer at Pewlett Hackard. Your first major task is a research project on employees of the corporation from the 1980s and 1990s. All that remain of the database of empl…
arxiv.org/abs/2505.10634v5
Over-reliance on language priors is a major cause of hallucinations in Large Vision-Language Models (LVLMs), often leading to outputs that are linguistically plausible but visually inconsistent. Recent studies have explored contrastive decoding as a...
arxiv.org/abs/2407.06438v3
We present SOLO, a single transformer for Scalable visiOn-Language mOdeling. Current large vision-language models (LVLMs) such as LLaVA mostly employ heterogeneous architectures that connect pre-trained visual encoders with large language models (LLM...
arxiv.org/abs/2312.07533v4
Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM with visual inputs, but lacks an in-depth study of the visual language p...
arxiv.org/abs/2503.13864v1
Detection of data races is one of the most important tasks for verifying the correctness of OpenMP parallel codes. Two main models of analysis tools have been proposed for detecting data races: dynamic analysis and static analysis. Dynamic analysis t...
arxiv.org/abs/2508.15503v4
Large Language Models (LLMs) are now ubiquitous in software engineering (SE) research and practice, yet their non-determinism, opaque training data, and rapidly evolving models threaten the reproducibility and replicability of empirical studies. We a...
openai.com/research/language-models-can-explain-neurons-in-language-models
Points: 688 | Comments: 477 | Author: mfiguiere
arxiv.org/abs/1909.12073v3
The synthetic control method (SCM) allows estimating the causal effect of an intervention in settings where panel data on a small number of treated and control units are available. We show that the existing SCM, as well as its extensions, can be easi...
arxiv.org/abs/2505.18995v1
This study presents FiLLM, a Filipino-optimized large language model, designed to enhance natural language processing (NLP) capabilities in the Filipino language. Built upon the SeaLLM-7B 2.5 model, FiLLM leverages Low-Rank Adaptation (LoRA) fine-tun...
github.com/nealcaren/social-data-analysis
Claude Code plugins for quantitative and qualitative sociological research (⭐ 14)
arxiv.org/abs/1807.01990v1
Capturing and labeling camera images in the real world is an expensive task, whereas synthesizing labeled images in a simulation environment is easy for collecting large-scale image data. However, learning from only synthetic images may not achieve t...