860 results for Evaluating · 0.078s

github.com/dannguyen/abbyy-finereader-ocr-senate

dannguyen/abbyy-finereader-ocr-senate

Evaluating the performance and accuracy of ABBYY FineReader's OCR on Senate Financial Disclosure scanned forms (⭐ 135)

Sponsored Partners
arxiv.org/abs/2601.07783v1

Affordable Data Collection System for UAVs Taxi Vibration Testing

Structural vibration testing plays a key role in aerospace engineering for evaluating dynamic behaviour, ensuring reliability and verifying structural integrity. These tests rely on accurate and robust data acquisition systems (DAQ) to capture high-q...

arxiv.org/abs/2510.12740v2

Hey, wait a minute: on at-issue sensitivity in Language Models

Evaluating the naturalness of dialogue in language models (LMs) is not trivial: notions of 'naturalness' vary, and scalable quantitative metrics remain limited. This study leverages the linguistic notion of 'at-issueness' to assess dialogue naturalne...

arxiv.org/abs/2310.03702v1

Robust Analysis of Auction Equilibria

Equilibria in auctions can be very difficult to analyze, beyond the symmetric environments where revenue equivalence renders the analysis straightforward. This paper takes a robust approach to evaluating the equilibria of auctions. Rather than identi...

arxiv.org/abs/1309.7266v1

Evaluating Link-Based Techniques for Detecting Fake Pharmacy Websites

Fake online pharmacies have become increasingly pervasive, constituting over 90% of online pharmacy websites. There is a need for fake website detection techniques capable of identifying fake online pharmacy websites with a high degree of accuracy. I...

arxiv.org/abs/2408.10395v1

Evaluating Image-Based Face and Eye Tracking with Event Cameras

Event Cameras, also known as Neuromorphic sensors, capture changes in local light intensity at the pixel level, producing asynchronously generated data termed ``events''. This distinct data format mitigates common issues observed in conventional came...

arxiv.org/abs/2306.11174v1

Evaluating Privacy Questions From Stack Overflow: Can ChatGPT Compete?

Stack Overflow and other similar forums are used commonly by developers to seek answers for their software development as well as privacy-related concerns. Recently, ChatGPT has been used as an alternative to generate code or produce responses to dev...

arxiv.org/abs/2105.00071v3

Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark

Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information. Progress towards models that do not exhibit this issue requires evaluation met...

arxiv.org/abs/2407.19897v1

BEExAI: Benchmark to Evaluate Explainable AI

Recent research in explainability has given rise to numerous post-hoc attribution methods aimed at enhancing our comprehension of the outputs of black-box machine learning models. However, evaluating the quality of explanations lacks a cohesive appro...

arxiv.org/abs/2503.17963v1

Won: Establishing Best Practices for Korean Financial NLP

In this work, we present the first open leaderboard for evaluating Korean large language models focused on finance. Operated for about eight weeks, the leaderboard evaluated 1,119 submissions on a closed benchmark covering five MCQA categories: finan...

en.wikipedia.org/wiki/Comparison

Comparison - Wikipedia

Comparison or comparing is the act of evaluating two or more things by determining the relevant, comparable characteristics of each thing, and then determining

github.com/TheBlackKazekage/FiredEmployees

TheBlackKazekage/FiredEmployees

Mr. Alfred is the founder of FooLand Constructions. He always maintains a ‘Black list’ of potential employees which can be fired at any moment without any prior notice. This company has N employees (including Mr. Alfred), and each employee is assigned a rank (1 <= rank <= N) at t…

github.com/edwanyoike/mr_alfred_and_employee_evaluation

edwanyoike/mr_alfred_and_employee_evaluation

Mr. Alfred is the founder of FooLand Constructions. He always maintains a ‘Black list’ of potential employees which can be fired at any moment without any prior notice. This company has N employees (including Mr. Alfred), and each employee is assigned a rank (1 <= rank <= N) at t…

arxiv.org/abs/1312.0341v1

The TTC 2013 Flowgraphs Case

This case for the Transformation Tool Contest 2013 is about evaluating the scope and usability of transformation languages and tools for a set of four tasks requiring very different capabilities. One task deals with typical model-to-model transforma...

arxiv.org/abs/1612.08913v1

Evaluating Marijuana-Related Tweets On Twitter

This paper studies marijuana-related tweets in social network Twitter. We collected more than 300,000 marijuana related tweets during November 2016 in our study. Our text-mining based algorithms and data analysis unveil some interesting patterns incl...

arxiv.org/abs/2601.14569v1

Social Caption: Evaluating Social Understanding in Multimodal Models

Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions. We introduce Social Caption, a framework grounded in interaction theory to evaluate social understanding abilities of MLLM...