860 results for Evaluating · 0.079s

arxiv.org/abs/2511.16035v2

Liars' Bench: Evaluating Lie Detectors for Language Models

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies L...

Sponsored Partners
arxiv.org/abs/2306.02077v1

Utilizing ChatGPT to Enhance Clinical Trial Enrollment

Clinical trials are a critical component of evaluating the effectiveness of new medical interventions and driving advancements in medical research. Therefore, timely enrollment of patients is crucial to prevent delays or premature termination of tria...

arxiv.org/abs/1305.5653v1

Geographica: A Benchmark for Geospatial RDF Stores

Geospatial extensions of SPARQL like GeoSPARQL and stSPARQL have recently been defined and corresponding geospatial RDF stores have been implemented. However, there is no widely used benchmark for evaluating geospatial RDF stores which takes into acc...

arxiv.org/abs/1904.12573v2

Venue Analytics: A Simple Alternative to Citation-Based Metrics

We present a method for automatically organizing and evaluating the quality of different publishing venues in Computer Science. Since this method only requires paper publication data as its input, we can demonstrate our method on a large portion of t...

arxiv.org/abs/2310.05060v2

DeVAn: Dense Video Annotation for Video-Language Models

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K YouTube video c...

arxiv.org/abs/2505.04110v1

Alpha Excel Benchmark

This study presents a novel benchmark for evaluating Large Language Models (LLMs) using challenges derived from the Financial Modeling World Cup (FMWC) Excel competitions. We introduce a methodology for converting 113 existing FMWC challenges into pr...

arxiv.org/abs/2406.08221v2

FAIL: Analyzing Software Failures from the News Using LLMs

Software failures inform engineering work, standards, regulations. For example, the Log4J vulnerability brought government and industry attention to evaluating and securing software supply chains. Accessing private engineering records is difficult, s...

arxiv.org/abs/1401.1918v1

Key Performance Indicators for QOS Assessment in TETRA Networks

Key Performance Indicators (KPIs) are widely used by GSM and UMTS carriers with the aim of evaluating the network performance and the Quality of Service (QoS) delivered to users. TETRA networks are basically designed to provide telecommunication serv...

www.bing.com/ck/a?!&&p=8cfefd327c49a70f09c4026189f3d76894bdf73f89f0353a5e03b6edeb82b33eJmltdHM9MTc3MjU4MjQwMA&ptn=3&ver=2&hsh=4&fclid=0eedaa02-ade4-6598-3e37-bd10ac0764d5&u=a1aHR0cHM6Ly9uZXdzLm1pdC5lZHUvMjAyNS9tYWtpbmctY2xlYW4tZW5lcmd5LWludmVzdG1lbnRzLW1vcmUtc3VjY2Vzc2Z1bC0xMjEy&ntb=1

Making clean energy investments more successful - MIT News

Dec 12, 2025 · New research emphasizes the importance of well-validated models and forecasting tools in evaluating choices for investments in clean energy technologies and policies by governments and …