1,230 results for Datasets · 0.078s

arxiv.org/abs/2402.14590v1

Scaling Up LLM Reviews for Google Ads Content Moderation

Large language models (LLMs) are powerful tools for content moderation, but their inference costs and latency make them prohibitive for casual use on large datasets, such as the Google Ads repository. This study proposes a method for scaling up LLM r...

Sponsored Partners
arxiv.org/abs/1512.04092v1

Stack Exchange Tagger

The goal of our project is to develop an accurate tagger for questions posted on Stack Exchange. Our problem is an instance of the more general problem of developing accurate classifiers for large scale text datasets. We are tackling the multilabel c...

github.com/Mo-Zeini/The-Impacts-of-Working-Remotely-and-in-an-Office

Mo-Zeini/The-Impacts-of-Working-Remotely-and-in-an-Office

In this project, we explore the reasons that lead to the negative impact of working remotely on mental and physical health and investigate whether employees are aware of the negative and the positive effects of working either from home or in an office using datasets and data anal…

arxiv.org/abs/2410.20335v2

Robust Universum Twin Support Vector Machine for Imbalanced Data

One of the major difficulties in machine learning methods is categorizing datasets that are imbalanced. This problem may lead to biased models, where the training process is dominated by the majority class, resulting in inadequate representation of t...

github.com/jlmelville/vizier

jlmelville/vizier

Visualization of 2D Datasets in R (⭐ 6)

www.bing.com/ck/a?!&&p=09345da27b35fada48ef27fc44484828d3d63b5879ac01613c0b159e7c7feba9JmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=356326bc-a1f2-65ed-2d48-31a8a033647d&u=a1aHR0cHM6Ly93d3cub2NlYW5vZ3JhcGh5LmNvbS8&ntb=1

Oceanography.com | An oceanographic learning and research …

Join a dynamic oceanographic community for research. Access real-time and historical datasets by region and explore the global ocean index.

arxiv.org/abs/2603.02638v1

The Vienna 4G/5G Drive-Test Dataset

Machine learning for mobile network analysis, planning, and optimization is often limited by the lack of large, comprehensive real-world datasets. This paper introduces the Vienna 4G/5G Drive-Test Dataset, a city-scale open dataset of georeferenced L...

www.bing.com/ck/a?!&&p=bd2632c26ebd845e794709fc10262bde00ec1ca9c3516abc01dab51ae0fab8a1JmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1f9ff355-f212-68c1-3cca-e441f3a4694f&u=a1aHR0cHM6Ly9yZXNlYXJjaC5nb29nbGUvcHVicy8&ntb=1

Publications – Google Research

Building on research that emphasizes the need for specialized datasets and model training tools, our study uses a scaffolded approach to understand the ideal model training and voice recording process.

arxiv.org/abs/1506.05101v1

Big Data Analytics in Bioinformatics: A Machine Learning Perspective

Bioinformatics research is characterized by voluminous and incremental datasets and complex data analytics methods. The machine learning methods used in bioinformatics are iterative and parallel. These methods can be scaled to handle big data using t...

arxiv.org/abs/2506.10217v3

Data-Centric Safety and Ethical Measures for Data and AI Governance

Datasets play a key role in imparting advanced capabilities to artificial intelligence (AI) foundation models that can be adapted to various downstream tasks. These downstream applications can introduce both beneficial and harmful capabilities -- res...

arxiv.org/abs/2407.01343v1

Coordination Failure in Cooperative Offline MARL

Offline multi-agent reinforcement learning (MARL) leverages static datasets of experience to learn optimal multi-agent control. However, learning from static data presents several unique challenges to overcome. In this paper, we focus on coordination...

www.bing.com/ck/a?!&&p=8b2555a962d56cdc8a9f796715d95cfe44a118e05d0263c9a4ea5fdc10549802JmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1913882f-5e4a-60da-1585-9f3b5f606164&u=a1aHR0cHM6Ly9tYXAuZ2VvLmFkbWluLmNoLz90b3BpYz1jYWRhc3RyZQ&ntb=1

Maps of Switzerland - Swiss Confederation - map.geo.admin.ch

Official Swiss platform for cadastral and geospatial data, providing access to maps, datasets, and services for various applications.

www.bing.com/ck/a?!&&p=41cb9ef2003341c752558c96554d8daa0fbbb47adbe8fe800aeb795ea4f8c2a5JmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1913882f-5e4a-60da-1585-9f3b5f606164&u=a1aHR0cHM6Ly9tYXAuZ2VvLmFkbWluLmNoL2luZGV4Lmh0bWw_c3dpc3NzZWFyY2g9MCww&ntb=1

Maps of Switzerland - Swiss Confederation - map.geo.admin.ch

Explore Switzerland's maps, geodata, and services with the Swiss Confederation's geoportal. Discover ski routes, weekly datasets, and more.

arxiv.org/abs/2509.12178v2

All that structure matches does not glitter

Generative models for materials, especially inorganic crystals, hold potential to transform the theoretical prediction of novel compounds and structures. Advancement in this field depends on robust benchmarks and minimal, information-rich datasets th...

github.com/uhh-lt/newsleak

uhh-lt/newsleak

Information extraction and interactive visualization of textual datasets for investigative data-driven journalism and eDiscovery (⭐ 60)

arxiv.org/abs/2306.03286v2

Survival Instinct in Offline Reinforcement Learning

We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe policies even when trained with "wrong" reward labels, such as those that...

arxiv.org/abs/2303.09093v3

GLEN: General-Purpose Event Detection for Thousands of Types

The progress of event extraction research has been hindered by the absence of wide-coverage, large-scale datasets. To make event extraction systems more accessible, we build a general-purpose event detection dataset GLEN, which covers 205K event ment...

arxiv.org/abs/2406.07539v2

BAKU: An Efficient Transformer for Multi-Task Policy Learning

Training generalist agents capable of solving diverse tasks is challenging, often requiring large datasets of expert demonstrations. This is particularly problematic in robotics, where each data point requires physical execution of actions in the rea...

arxiv.org/abs/2406.10531v1

PIG: Prompt Images Guidance for Night-Time Scene Parsing

Night-time scene parsing aims to extract pixel-level semantic information in night images, aiding downstream tasks in understanding scene object distribution. Due to limited labeled night image datasets, unsupervised domain adaptation (UDA) has becom...

github.com/aliyarahman/theres-a-job-for-that

aliyarahman/theres-a-job-for-that

(Ideation materials/Flask) White House Data Jam rapid prototype. Simple, engaging user interface to assist job seekers in navigating existing national datasets and APIs. (⭐ 3)

arxiv.org/abs/2404.12712v1

uTRAND: Unsupervised Anomaly Detection in Traffic Trajectories

Deep learning-based approaches have achieved significant improvements on public video anomaly datasets, but often do not perform well in real-world applications. This paper addresses two issues: the lack of labeled data and the difficulty of explaini...