860 results for Evaluating · 0.810s

Sponsored Partners
www.desmos.com/scientific

Desmos | Scientific Calculator

A beautiful, free online scientific calculator with advanced features for evaluating percentages, fractions, exponential functions, logarithms, trigonometry, statistics, and more.

arxiv.org/abs/1612.03353v2

FOCA: A Methodology for Ontology Evaluation

Modeling an ontology is a hard and time-consuming task. Although methodologies are useful for ontologists to create good ontologies, they do not help with the task of evaluating the quality of the ontology to be reused. For these reasons, it is imper...

arxiv.org/abs/2108.13703v1

Evaluating the Robustness of Off-Policy Evaluation

Off-policy Evaluation (OPE), or offline evaluation in general, evaluates the performance of hypothetical policies leveraging only offline log data. It is particularly useful in applications where the online interaction involves high stakes and expens...

arxiv.org/abs/2309.02133v1

Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion

Foreign accent conversion (FAC) is a special application of voice conversion (VC) which aims to convert the accented speech of a non-native speaker to a native-sounding speech with the same speaker identity. FAC is difficult since the native speech f...

measure-wellbeing.org/wellbeing-explained

What is wellbeing, and what matters? – Evaluating wellbeing

One way of understanding wellbeing is how well people are able to flourish – whether they feel positive emotions, can function well in society, can respond to challenges and make meaning in their lives – …

www.monster.com/career-advice/job-search?msockid=16a7c07c3d5b6ad10b83d7693c916bd4

Job Search Archives - Monster

Learn how to get noticed by recruiters, along with tips and tricks for networking, gathering references, organizing your search, evaluating job offers, and more from the career experts at Monster.

github.com/microsoft/contractor

microsoft/contractor

A repository that implements a agentic system capable of evaluating security and quotization contracts based on a wide range of data sources. (⭐ 15)

arxiv.org/abs/cs/0205008v1

Improved Bicriteria Existence Theorems for Scheduling

Two common objectives for evaluating a schedule are the makespan, or schedule length, and the average completion time. This short note gives improved bounds on the existence of schedules that simultaneously optimize both criteria. In particular, fo...

github.com/VietnamAIHub/Vietnamese_LLMs

VietnamAIHub/Vietnamese_LLMs

Dự án bao gồm: 1. Xây dựng bộ dữ Instructions Vietnamese (chất lượng, nhiều, và đa dạng). 2.LLM Training, Finetuning, Evaluating & Testing trên Open-source mô hình ngôn ngữ: Bloomz,T5, UL2, LLaMA (1&2), OpenLLaMA, GPT-J pythia etc. 3. Ứng dụng và Giao diện Người dùng (UI) (⭐ 278)

arxiv.org/abs/2409.07641v1

SimulBench: Evaluating Language Models with Creative Simulation Tasks

We introduce SimulBench, a benchmark designed to evaluate large language models (LLMs) across a diverse collection of creative simulation scenarios, such as acting as a Linux terminal or playing text games with users. While these simulation tasks ser...

measure-wellbeing.org/wellbeing-explained

What is wellbeing, and what matters? – Evaluating wellbeing

Wellbeing is responsive and dynamic – it’s not static but changes as different factors and aspects of people’s lives change. This dynamic nature of wellbeing was the basis of New Economic Foundation …

measure-wellbeing.org/wellbeing-explained

What is wellbeing, and what matters? – Evaluating wellbeing

Wellbeing is ‘how we are doing’ as individuals, communities and as a nation, and how sustainable this is for the future. It isn’t just about how things look from the outside, it’s about how we feel in ourselves.

arxiv.org/abs/2510.00172v1

DRBench: A Realistic Benchmark for Enterprise Deep Research

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (...

arxiv.org/abs/2111.07201v2

Evaluating the effectiveness of Phishing Reports on Twitter

Phishing attacks are an increasingly potent web-based threat, with nearly 1.5 million websites created on a monthly basis. In this work, we present the first study towards identifying such attacks through phishing reports shared by users on Twitter....

arxiv.org/abs/1906.08600v1

CBC Approach for Evaluating Potential SaaS on the Cloud

The cloud computing is evolving as a key computing platform for sharing resources like infrastructure, platform, software etc. This has proven to be an essential requirement for extending many existing applications. Software as a service (SaaS) is re...

arxiv.org/abs/1912.07364v1

Evaluating one-shot tournament predictions

We introduce the Tournament Rank Probability Score (TRPS) as a measure to evaluate and compare pre-tournament predictions, where predictions of the full tournament results are required to be available before the tournament begins. The TRPS handles pa...

brilliant.org/wiki/integration

Integration | Brilliant Math & Science Wiki

Integration is the process of evaluating integrals. It is one of the two central ideas of calculus and is the inverse of the other central idea of calculus, differentiation.

www.desmos.com/scientific

Desmos | Scientific Calculator

A beautiful, free online scientific calculator with advanced features for evaluating percentages, fractions, exponential functions, logarithms, trigonometry, statistics, and more.