860 results for Evaluating · 0.118s

github.com/dannguyen/abbyy-finereader-ocr-senate

dannguyen/abbyy-finereader-ocr-senate

Evaluating the performance and accuracy of ABBYY FineReader's OCR on Senate Financial Disclosure scanned forms (⭐ 135)

Sponsored Partners
arxiv.org/abs/2601.07783v1

Affordable Data Collection System for UAVs Taxi Vibration Testing

Structural vibration testing plays a key role in aerospace engineering for evaluating dynamic behaviour, ensuring reliability and verifying structural integrity. These tests rely on accurate and robust data acquisition systems (DAQ) to capture high-q...

arxiv.org/abs/2510.12740v2

Hey, wait a minute: on at-issue sensitivity in Language Models

Evaluating the naturalness of dialogue in language models (LMs) is not trivial: notions of 'naturalness' vary, and scalable quantitative metrics remain limited. This study leverages the linguistic notion of 'at-issueness' to assess dialogue naturalne...

arxiv.org/abs/2310.03702v1

Robust Analysis of Auction Equilibria

Equilibria in auctions can be very difficult to analyze, beyond the symmetric environments where revenue equivalence renders the analysis straightforward. This paper takes a robust approach to evaluating the equilibria of auctions. Rather than identi...

arxiv.org/abs/1309.7266v1

Evaluating Link-Based Techniques for Detecting Fake Pharmacy Websites

Fake online pharmacies have become increasingly pervasive, constituting over 90% of online pharmacy websites. There is a need for fake website detection techniques capable of identifying fake online pharmacy websites with a high degree of accuracy. I...

arxiv.org/abs/2408.10395v1

Evaluating Image-Based Face and Eye Tracking with Event Cameras

Event Cameras, also known as Neuromorphic sensors, capture changes in local light intensity at the pixel level, producing asynchronously generated data termed ``events''. This distinct data format mitigates common issues observed in conventional came...

arxiv.org/abs/2306.11174v1

Evaluating Privacy Questions From Stack Overflow: Can ChatGPT Compete?

Stack Overflow and other similar forums are used commonly by developers to seek answers for their software development as well as privacy-related concerns. Recently, ChatGPT has been used as an alternative to generate code or produce responses to dev...

arxiv.org/abs/2105.00071v3

Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark

Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information. Progress towards models that do not exhibit this issue requires evaluation met...

arxiv.org/abs/2407.19897v1

BEExAI: Benchmark to Evaluate Explainable AI

Recent research in explainability has given rise to numerous post-hoc attribution methods aimed at enhancing our comprehension of the outputs of black-box machine learning models. However, evaluating the quality of explanations lacks a cohesive appro...

arxiv.org/abs/2503.17963v1

Won: Establishing Best Practices for Korean Financial NLP

In this work, we present the first open leaderboard for evaluating Korean large language models focused on finance. Operated for about eight weeks, the leaderboard evaluated 1,119 submissions on a closed benchmark covering five MCQA categories: finan...