Information retrieval (IR) benchmarks typically follow the Cranfield paradigm, relying on static and predefined corpora. However, temporal changes in technical corpora, such as API deprecations and code reorganizations, can render existing benchmarks...
In this paper, we investigate the use of large language models (LLMs) like ChatGPT for document-grounded response generation in the context of information-seeking dialogues. For evaluation, we use the MultiDoc2Dial corpus of task-oriented dialogues i...
Conversational question answering aims to provide natural-language answers to users in information-seeking conversations. Existing conversational QA benchmarks compare models with pre-collected human-human conversations, using ground-truth answers pr...
Implementing enterprise process automation often requires significant technical expertise and engineering effort. It would be beneficial for non-technical users to be able to describe a business process in natural language and have an intelligent sys...
The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy issues for user...
A challenge that data analysts face is building a data analysis that is useful for a given consumer. Previously, we defined a set of principles for describing data analyses that can be used to create a data analysis and to characterize the variation...
After sifting the sources and evaluating them, the life of John the son of Zebedee may be summarized in the following sequence. He was a convert of John the Baptist and spent some time with the …
Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for a more robust multi-prompt evaluation is challenging, limiting its ado...
Adopting a systematic approach to identifying risk factors, assessing severity, and evaluating probability provides a solid foundation for effective risk ranking. In summary, grasping the characteristics of risks …
Risk Ranking and Filtering works by breaking down Definition: SYSTEM is the overall risk into risk components and evaluating those subject of a risk components and their individual contributions to …
Evaluating the effectiveness and benefits of driver assistance systems is essential for improving the system performance. In this paper, we propose an efficient evaluation method for a semi-autonomous lane departure correction system. To achieve this...
The AWS Well-Architected Tool , available at no cost in the AWS Management Console , provides a mechanism for regularly evaluating workloads, identifying high-risk issues, and recording …
This study investigates the relationship between sugarcane yield and cane height derived under different water and nitrogen conditions from pre-harvest Digital Surface Model (DSM) obtained via Unmanned Aerial Vehicle (UAV) flights over a sugarcane te...
Evaluating large language models (LLMs) as judges is increasingly critical for building scalable and trustworthy evaluation pipelines. We present ScalingEval, a large-scale benchmarking study that systematically compares 36 LLMs, including GPT, Gemin...
The offset method for solving word analogies has become a standard evaluation tool for vector-space semantic models: it is considered desirable for a space to represent semantic relations as consistent vector offsets. We show that the method's relian...
Randomized sampling based algorithms are widely used in robot motion planning due to the problem's intractability, and are experimentally effective on a wide range of problem instances. Most variants do not sample uniformly at random, and instead bia...