860 results for Evaluating · 0.093s

arxiv.org/abs/2504.03889v4

Identifying and Evaluating Inactive Heads in Pretrained LLMs

Attention is foundational to large language models (LLMs), enabling different heads to have diverse focus on relevant input tokens. However, learned behaviors like attention sinks, where the first token receives the most attention despite limited sem...

github.com/cheesesashimi/theywhiteboardedme

cheesesashimi/theywhiteboardedme

List of companies that use whiteboarding, live-coding or other high-pressure interview tactics for evaluating software engineering candidates. (⭐ 118)

Sponsored Partners
arxiv.org/abs/2311.05076v4

Evaluating diversion and treatment policies for opioid use disorder

The United States (US) opioid crisis contributed to 81,806 fatalities in 2022. It has strained hospitals, treatment facilities, and law enforcement agencies due to the enormous resources and procedures needed to respond to the crisis. As a result, ma...

arxiv.org/abs/2310.12971v1

CLAIR: Evaluating Image Captions with Large Language Models

The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, capt...

arxiv.org/abs/2409.12962v2

CLAIR-A: Leveraging Large Language Models to Judge Audio Captions

The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex task that requires considering diverse factors, among them, auditory sce...

arxiv.org/abs/2010.07386v2

Valentine: Evaluating Matching Techniques for Dataset Discovery

Data scientists today search large data lakes to discover and integrate datasets. In order to bring together disparate data sources, dataset discovery methods rely on some form of schema matching: the process of establishing correspondences between d...

arxiv.org/abs/2512.10675v2

Evaluating Gemini Robotics Policies in a Veo World Simulator

Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable and gener...

arxiv.org/abs/2511.02120v1

Evaluating Factor Contributions for Sold Homes

We evaluate the contributions of ten intrinsic and extrinsic factors, including ESG (environmental, social, and governance) factors readily available from website data to individual home sale prices using a P-spline generalized additive model (GAM)....