COCO api evaluation for subset of classes
Tags: python, subset, mscoco, pycocotools | Score: 10
Tags: python, subset, mscoco, pycocotools | Score: 10
Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark saturation, lack of...
Understanding and identifying musical shape plays an important role in music education and performance assessment. To simplify the otherwise time- and cost-intensive musical shape evaluation, in this paper we explore how artificial intelligence (AI)...
We present QGen Studio: an adaptive question-answer generation, training, and evaluation platform. QGen Studio enables users to leverage large language models (LLMs) to create custom question-answer datasets and fine-tune models on this synthetic dat...
Shared challenges provide a venue for comparing systems trained on common data using a standardized evaluation, and they also provide an invaluable resource for researchers when the data and evaluation results are publicly released. The Blizzard Chal...
Huber loss, its asymmetric variants and their associated functionals (here named Huber functionals) are studied in the context of point forecasting and forecast evaluation. The Huber functional of a distribution is the set of minimizers of the expect...
In some research evaluation systems, credit awarded to an article depends on the number of co-authors on the article with total credit to the article increasing with the number of co-authors. There are many examples of such evaluation systems (e.g.,...
The recent endeavors of the research community to unite efforts on the design and evaluation of congestion control algorithms have created a growing collection of congestion control schemes called Pantheon. However, the virtual network emulator that...
The taste electroencephalogram (EEG) evoked by the taste stimulation can reflect different brain patterns and be used in applications such as sensory evaluation of food. However, considering the computational cost and efficiency, EEG data with many c...
For humans, taste is essential for perceiving food's nutrient content or harmful components. The current sensory evaluation of taste mainly relies on artificial sensory evaluation and electronic tongue, but the former has strong subjectivity and poor...
Are you looking for an ADHD evaluation and follow-up in Montréal, Laval, Pierrefonds, or on the South Shore? At Union MD Clinic, our mental-health specialists offer complete assessments for Attention …
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist to...
Ensuring the long-term sustainability of recommender systems (RS) emerges as a crucial issue. Traditional offline evaluation methods for RS typically focus on immediate user feedback, such as clicks, but they often neglect the long-term impact of con...
Multimodal generative AI usually involves generating image or text responses given inputs in another modality. The evaluation of image-text relevancy is essential for measuring response quality or ranking candidate responses. In particular, binary re...
Visual quality evaluation is one of the challenging basic problems in image processing. It also plays a central role in the shaping, implementation, optimization, and testing of many methods. The existing image quality assessment methods focused on i...
In this work, we presented a study regarding two important aspects of evolving feature-based game evaluation functions: the choice of genome representation and the choice of opponent used to test the model. We compared three representations. One simp...
Effective evaluation of web data record extraction methods is crucial, yet hampered by static, domain-specific benchmarks and opaque scoring practices. This makes fair comparison between traditional algorithmic techniques, which rely on structural he...
Psychological evaluations are usually requested by a patient's physician or therapist and are designed to assist the referring provider with respect to diagnostic clarification and treatment planning. …
In this paper we extend our techniques, developed in a previous paper (Du, etc, JHEP 05(2016)086) for direct evaluation of arbitrary $n$-point tree-level MHV amplitudes in 4d Yang-Mills and gravity theory using the Cachazo-He-Yuan (CHY) formalism, to...
The results of a macro-scale experimental study performed on a hardened class G cement paste [Ghabezloo et al. (2008) Cem. Con. Res. (38) 1424-1437] are used in association with the micromechanics modelling and homogenization technique for evaluation...