1,230 results for Datasets · 1.414s

arxiv.org/abs/2203.12637v1

Asynchronous Collaborative Learning Across Data Silos

Machine learning algorithms can perform well when trained on large datasets. While large organisations often have considerable data assets, it can be difficult for these assets to be unified in a manner that makes training possible. Data is very ofte...

Sponsored Partners
arxiv.org/abs/2505.12568v1

Enriching Patent Claim Generation with European Patent Dataset

Drafting patent claims is time-intensive, costly, and requires professional skill. Therefore, researchers have investigated large language models (LLMs) to assist inventors in writing claims. However, existing work has largely relied on datasets from...

www.reddit.com/r/doctorsUK/comments/1atejho/why_does_111_exist/

Why does 111 exist ?

Datasets show most providers achieve less than 10% end to end call closure. Patients wait hours only to be signposted elsewhere after a lengthy history taking exercise. Either their clinicians are not...

github.com/Nicheen/CuttingTools

Nicheen/CuttingTools

The datasets used in the thesis "Cutting Tool Container Inspection: Stereo vision and monocular artificial intelligence depth estimation at Sandvik Coromant" are presented in this repository. (⭐ 1)

arxiv.org/abs/2309.04951v2

Multi-document Summarization: A Comparative Evaluation

This paper is aimed at evaluating state-of-the-art models for Multi-document Summarization (MDS) on different types of datasets in various domains and investigating the limitations of existing models to determine future research directions. To addres...

arxiv.org/abs/2112.00050v1

Pattern-Aware Data Augmentation for LiDAR 3D Object Detection

Autonomous driving datasets are often skewed and in particular, lack training data for objects at farther distances from the ego vehicle. The imbalance of data causes a performance degradation as the distance of the detected objects increases. In thi...

github.com/nlextract/NLExtract

nlextract/NLExtract

Convert (ETL) and visualize free Dutch geo-datasets. (⭐ 167)

arxiv.org/abs/2512.14439v1

VICTOR: Dataset Copyright Auditing in Video Recognition Systems

Video recognition systems are increasingly being deployed in daily life, such as content recommendation and security monitoring. To enhance video recognition development, many institutions have released high-quality public datasets with open-source l...

github.com/usnistgov/SimulatedRadarWaveformGenerator

usnistgov/SimulatedRadarWaveformGenerator

A software tool that generates simulated radar signals and creates RF datasets for developing and testing machine/deep learning detection algorithms. (⭐ 106)

arxiv.org/abs/2307.16795v1

Structural Transfer Learning in NL-to-Bash Semantic Parsers

Large-scale pre-training has made progress in many fields of natural language processing, though little is understood about the design of pre-training datasets. We propose a methodology for obtaining a quantitative understanding of structural overlap...

github.com/ciads-ut/transfer-learning-ner

ciads-ut/transfer-learning-ner

Code, datasets, and results for "Transfer Learning for Entity Recognition of Novel Classes," Rodriguez, Caldwell, and Liu (COLING, 2018). (⭐ 36)

arxiv.org/abs/1403.3495v1

Analyzing Large Biological Datasets with an Improved Algorithm for MIC

A computational framework utilizes the traditional similarity measures for mining the significant relationships in biological annotations is recently proposed by Tatiana V. Karpinets et al. [2]. In this paper, an improved approximation algorithm for...

arxiv.org/abs/2403.03881v4

Unlocking Dataset Distillation with Diffusion Models

Dataset distillation seeks to condense datasets into smaller but highly representative synthetic samples. While diffusion models now lead all generative benchmarks, current distillation methods avoid them and rely instead on GANs or autoencoders, or,...

arxiv.org/abs/2506.05334v2

Search Arena: Analyzing Search-Augmented LLMs

Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in scope, of...

arxiv.org/abs/1906.04082v1

Char-RNN for Word Stress Detection in East Slavic Languages

We explore how well a sequence labeling approach, namely, recurrent neural network, is suited for the task of resource-poor and POS tagging free word stress detection in the Russian, Ukranian, Belarusian languages. We present new datasets, annotated...

arxiv.org/abs/2307.11661v2

Enhancing CLIP with GPT-4: Harnessing Visual Descriptions as Prompts

Contrastive pretrained large Vision-Language Models (VLMs) like CLIP have revolutionized visual representation learning by providing good performance on downstream datasets. VLMs are 0-shot adapted to a downstream dataset by designing prompts that ar...

github.com/andresilvapimentel/endocrine-disruption-explainer

andresilvapimentel/endocrine-disruption-explainer

Endocrine Disruption Explainer is a code to generate structural alerts of endocrine disruption of chemcial compounds using Local Interpretable Model-Agnostic Explanations (LIME) of machine learning models from TOX-21, EDC, and EDKB-FDA datasets. (⭐ 3)

arxiv.org/abs/2206.03693v3

Autoregressive Perturbations for Data Poisoning

The prevalence of data scraping from social media as a means to obtain datasets has led to growing concerns regarding unauthorized use of data. Data poisoning attacks have been proposed as a bulwark against scraping, as they make data "unlearnable" b...

arxiv.org/abs/2406.09173v3

Potion: Towards Poison Unlearning

Adversarial attacks by malicious actors on machine learning systems, such as introducing poison triggers into training datasets, pose significant risks. The challenge in resolving such an attack arises in practice when only a subset of the poisoned d...