1,230 results for Datasets (0.711 seconds)

arxiv.org/abs/2207.01760v1

GP22: A Car Styling Dataset for Automotive Designers

An automated design data archiving could reduce the time wasted by designers from working creatively and effectively. Though many datasets on classifying, detecting, and instance segmenting on car exterior exist, these large datasets are not relevant...

arxiv.org/abs/2310.15239v1

CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks

Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks. However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios t...

arxiv.org/abs/2406.16966v1

Mitigating Noisy Supervision Using Synthetic Samples with Soft Labels

Noisy labels are ubiquitous in real-world datasets, especially in the large-scale ones derived from crowdsourcing and web searching. It is challenging to train deep neural networks with noisy datasets since the networks are prone to overfitting the n...

arxiv.org/abs/2506.10908v3

Probably Approximately Correct Labels

Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunity to automatically label datasets and save costs. Unfortunately, these...

arxiv.org/abs/2010.07061v3

GiantMIDI-Piano: A large-scale MIDI dataset for classical piano music

Symbolic music datasets are important for music information retrieval and musical analysis. However, there is a lack of large-scale symbolic datasets for classical piano music. In this article, we create a GiantMIDI-Piano (GP) dataset containing 38,7...

arxiv.org/abs/2108.10665v1

Sharing Practices for Datasets Related to Accessibility and Aging

Datasets sourced from people with disabilities and older adults play an important role in innovation, benchmarking, and mitigating bias for both assistive and inclusive AI-infused applications. However, they are scarce. We conduct a systematic review...

arxiv.org/abs/2105.01475v1

Insights on the V3C2 Dataset

For research results to be comparable, it is important to have common datasets for experimentation and evaluation. The size of such datasets, however, can be an obstacle to their use. The Vimeo Creative Commons Collection (V3C) is a video dataset des...

arxiv.org/abs/1506.00022v1

Graph Watermarks

From network topologies to online social networks, many of today's most sensitive datasets are captured in large graphs. A significant challenge facing owners of these datasets is how to share sensitive graphs with collaborators and authorized users,...

arxiv.org/abs/2409.06892v1

Formative Study for AI-assisted Data Visualization

This formative study investigates the impact of data quality on AI-assisted data visualizations, focusing on how uncleaned datasets influence the outcomes of these tools. By generating visualizations from datasets with inherent quality issues, the re...

github.com/google-research-datasets/scin

google-research-datasets/scin

The SCIN dataset contains 10,000+ images of dermatology conditions, crowdsourced with informed consent from US internet users. Contributions include self-reported demographic and symptom information and dermatologist labels. The dataset also contains estimated Fitzpatrick skin ty…