1,230 results for Datasets · 0.105s

arxiv.org/abs/2309.11827v1

The Impact of Silence on Speech Anti-Spoofing

The current speech anti-spoofing countermeasures (CMs) show excellent performance on specific datasets. However, removing the silence of test speech through Voice Activity Detection (VAD) can severely degrade performance. In this paper, the impact of...

Sponsored Partners
github.com/langfuse/langfuse

langfuse/langfuse

? Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. ?YC W23 (⭐ 22726)

www.bing.com/ck/a?!&&p=cb78798c4d1f40b0ba54ab664a7e3dc9ff61297168d284d966bbfc75da68361dJmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1ea15943-9543-6b05-3d93-4e5794126a73&u=a1aHR0cHM6Ly9hcmFuLWxhYi5jb20vc29mdHdhcmUvc2luZ2xlci8&ntb=1

SingleR | Aran Lab @ Technion

Here, we present SingleR, a novel computational method for unbiased cell type recognition of scRNA-seq. SingleR leverages reference transcriptomic datasets of pure cell types to infer the cell of origin …

www.bing.com/ck/a?!&&p=52be30221e4c03ab7343631db38d211a0b7f9e53fbdd99bbe55fa9bf4e58699eJmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1ea15943-9543-6b05-3d93-4e5794126a73&u=a1aHR0cHM6Ly9jcmFuLmRldi9TaW5nbGVS&ntb=1

SingleR: Reference-Based Single-Cell RNA-Seq Annotation

SingleR: Reference-Based Single-Cell RNA-Seq Annotation Performs unbiased cell type recognition from single-cell RNA sequencing data, by leveraging reference transcriptomic datasets of pure cell …

www.bing.com/ck/a?!&&p=50a5f39fdf091a4ebc2276a6b7939af512823af7f84ee7f23b2d4f512555626cJmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=1ea15943-9543-6b05-3d93-4e5794126a73&u=a1aHR0cHM6Ly9naXRodWIuY29tL2R2aXJhcmFuL1NpbmdsZVI&ntb=1

GitHub - dviraran/SingleR: SingleR: Single-cell RNA-seq cell types ...

Here, we present SingleR, a novel computational method for unbiased cell type recognition of scRNA-seq. SingleR leverages reference transcriptomic datasets of pure cell types to infer the cell of origin …

www.bing.com/ck/a?!&&p=f1ba49f3efb39351b0847be49e48b168a0b90e04e2936e2da23ff0fe295aa1a1JmltdHM9MTc3Mjc1NTIwMA&ptn=3&ver=2&hsh=4&fclid=0f71add3-4ca2-6210-189c-bac74d9263d1&u=a1aHR0cHM6Ly9zdGFja292ZXJmbG93LmNvbS9xdWVzdGlvbnMvMjI1ODMzOTEvcGVhay1zaWduYWwtZGV0ZWN0aW9uLWluLXJlYWx0aW1lLXRpbWVzZXJpZXMtZGF0YQ&ntb=1

algorithm - Peak signal detection in realtime timeseries data - Stack ...

Robust peak detection algorithm (using z-scores) I came up with an algorithm that works very well for these types of datasets. It is based on the principle of dispersion: if a new datapoint is a given x …

github.com/InteragencyEcologicalProgram/zooper

InteragencyEcologicalProgram/zooper

R package to download and integrate zooplankton datasets from the Sacramento San Joaquin Delta (⭐ 4)

arxiv.org/abs/1211.6014v1

Exploring the Mobility of Mobile Phone Users

Mobile phone datasets allow for the analysis of human behavior on an unprecedented scale. The social network, temporal dynamics and mobile behavior of mobile phone users have often been analyzed independently from each other using mobile phone datase...

github.com/GeoSprocket/ssudan-eco

GeoSprocket/ssudan-eco

Repository for datasets and tools related to biodiversity assessment in South Sudan (⭐ 5)

arxiv.org/abs/2512.02192v1

Story2MIDI: Emotionally Aligned Music Generation from Text

In this paper, we introduce Story2MIDI, a sequence-to-sequence Transformer-based model for generating emotion-aligned music from a given piece of text. To develop this model, we construct the Story2MIDI dataset by merging existing datasets for sentim...

arxiv.org/abs/2106.15434v1

Zoo-Tuning: Adaptive Transfer from a Zoo of Models

With the development of deep networks on various large-scale datasets, a large zoo of pretrained models are available. When transferring from a model zoo, applying classic single-model based transfer learning methods to each source model suffers from...

arxiv.org/abs/2505.21979v3

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly design...

arxiv.org/abs/2402.17517v1

Label-Noise Robust Diffusion Models

Conditional diffusion models have shown remarkable performance in various generative tasks, but training them requires large-scale datasets that often contain noise in conditional inputs, a.k.a. noisy labels. This noise leads to condition mismatch an...

arxiv.org/abs/2307.16686v1

Guiding Image Captioning Models Toward More Specific Captions

Image captioning is conventionally formulated as the task of generating captions for images that match the distribution of reference image-caption pairs. However, reference captions in standard captioning datasets are short and may not uniquely ident...

arxiv.org/abs/2208.08091v1

In-vehicle alertness monitoring for older adults

Alertness monitoring in the context of driving improves safety and saves lives. Computer vision based alertness monitoring is an active area of research. However, the algorithms and datasets that exist for alertness monitoring are primarily aimed at...

arxiv.org/abs/2309.11381v1

Studying Lobby Influence in the European Parliament

We present a method based on natural language processing (NLP), for studying the influence of interest groups (lobbies) in the law-making process in the European Parliament (EP). We collect and analyze novel datasets of lobbies' position papers and s...

www.bing.com/ck/a?!&&p=af573f3ab3d11a89e49b67957fd012800a8c1997ec0b0dae7bfeb74c5da63a39JmltdHM9MTc3MjY2ODgwMA&ptn=3&ver=2&hsh=4&fclid=255dbbf2-f75a-6824-39ae-ace6f6e46922&u=a1aHR0cHM6Ly9sZWFybi5taWNyb3NvZnQuY29tL2VuLXVzL2Fuc3dlcnMvcXVlc3Rpb25zLzU4MDkxNjAvZGF0YS1icmlja3MtZGF0YS1pbmdlc3Rpb24tdGVjaG5pcXVlcw&ntb=1

data bricks data ingestion techniques - Microsoft Q&A

16 hours ago · Hi Expert, needs following information Overall Purpose The board shows a Data Cycle / Data Strategy POC used to: Identify the right datasets Secure access to data sources Ingest and …

arxiv.org/abs/1810.11067v1

Teaching Syntax by Adversarial Distraction

Existing entailment datasets mainly pose problems which can be answered without attention to grammar or word order. Learning syntax requires comparing examples where different grammar and word order change the desired classification. We introduce sev...

arxiv.org/abs/2506.02071v1

AI Data Development: A Scorecard for the System Card Framework

Artificial intelligence has transformed numerous industries, from healthcare to finance, enhancing decision-making through automated systems. However, the reliability of these systems is mainly dependent on the quality of the underlying datasets, rai...

arxiv.org/abs/2103.07191v2

Are NLP Models really able to Solve Simple Math Word Problems?

The problem of designing NLP solvers for math word problems (MWP) has seen sustained research activity and steady gains in the test accuracy. Since existing solvers achieve high performance on the benchmark datasets for elementary level MWPs containi...

github.com/kastle-lab/KWG-LandUse

kastle-lab/KWG-LandUse

This repository contains the documentation, datasets, schema's, and code for Michael McCain's Knowledge Graph Land Use project. (⭐ 2)