arxiv.org/abs/2410.18678v1
This paper introduces Ali-AUG, a novel single-step diffusion model for efficient labeled data augmentation in industrial applications. Our method addresses the challenge of limited labeled data by generating synthetic, labeled images with precise fea...
arxiv.org/abs/1610.01670v1
We describe the Gun Violence Database (GVDB), a large and growing database of gun violence incidents in the United States. The GVDB is built from the detailed information found in local news reports about gun violence, and is constructed via a large-...
github.com/optery/optery-data-brokers-directory
Welcome to Optery’s open-source directory of data brokers and opt-out information, the largest of its kind. (⭐ 45)
github.com/adactio/TheSession-data
Data dumps from thesession.org (⭐ 81)
www.bing.com/ck/a?!&&p=89e13acd8fd4466e5bb3278c3e58d008cef073d52b8074587d332006f36fab5eJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=0b7deb60-2cf3-60f8-146e-fc712d196168&u=a1aHR0cHM6Ly93d3cubXNzcWx0aXBzLmNvbS9zcWxzZXJ2ZXJ0aXAvODAxNi9jb25maWd1cmUtbWljcm9zb2Z0LWZhYnJpYy1kYXRhYmFzZS1taXJyb3JpbmctZm9yLXNub3dmbGFrZS8&ntb=1
Jun 27, 2024 · Any database in Snowflake can be mirrored (also cloned databases), and thereâs a sample database in Snowflake to get you started. The following are prerequisites to begin with â¦
arxiv.org/abs/2009.04459v2
In this paper we present a new dataset, with musical excepts from the three main ethnic groups in Singapore: Chinese, Malay and Indian (both Hindi and Tamil). We use this new dataset to train different classification models to distinguish the origin...
www.bing.com/ck/a?!&&p=b73d9e762ed02d816db3dca386d8923826ff7fe2fb57a58bd161ab0e85a9fd7cJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=3c44b1d7-77b9-61fd-0885-a6c6765d6064&u=a1aHR0cHM6Ly9zdGFja292ZXJmbG93LmNvbS9xdWVzdGlvbnMvMTk4Mjg4MjIvaG93LWRvLWktY2hlY2staWYtYS1wYW5kYXMtZGF0YWZyYW1lLWlzLWVtcHR5&ntb=1
Nov 7, 2013 · How do I check if a pandas DataFrame is empty? I'd like to print some message in the terminal if the DataFrame is empty.
github.com/statsbomb/open-data
Free football data from StatsBomb (⭐ 3029)
arxiv.org/abs/2102.04462v3
The count-min sketch (CMS) is a time and memory efficient randomized data structure that provides estimates of tokens' frequencies in a data stream of tokens, i.e. point queries, based on random hashed data. A learning-augmented version of the CMS, r...
arxiv.org/abs/2305.15723v1
In this paper, we study the setting in which data owners train machine learning models collaboratively under a privacy notion called joint differential privacy [Kearns et al., 2018]. In this setting, the model trained for each data owner $j$ uses $j$...
arxiv.org/abs/2106.05468v2
Vertical Federated Learning (VFL) refers to the collaborative training of a model on a dataset where the features of the dataset are split among multiple data owners, while label information is owned by a single data owner. In this paper, we propose...
arxiv.org/abs/2205.01791v1
We present TartanDrive, a large scale dataset for learning dynamics models for off-road driving. We collected a dataset of roughly 200,000 off-road driving interactions on a modified Yamaha Viking ATV with seven unique sensing modalities in diverse t...
arxiv.org/abs/1810.06936v2
Data-driven algorithms have surpassed traditional techniques in almost every aspect in robotic vision problems. Such algorithms need vast amounts of quality data to be able to work properly after their training process. Gathering and annotating that...
arxiv.org/abs/2412.06288v2
The surging demand for AI has led to a rapid expansion of energy-intensive data centers, impacting the environment through escalating carbon emissions and water consumption. While significant attention has been paid to data centers' growing environme...
arxiv.org/abs/2002.07397v1
Open-domain retrieval-based dialogue systems require a considerable amount of training data to learn their parameters. However, in practice, the negative samples of training data are usually selected from an unannotated conversation data set at rando...
arxiv.org/abs/2105.14284v3
Management of invasive species and pathogens requires information about the traffic of potential vectors. Such information is often taken from vector traffic models fitted to survey data. Here, user-specific data collected via mobile apps offer new o...
arxiv.org/abs/1511.07772v1
The Workshop on Nuclear Data Needs and Capabilities for Applications (NDNCA) was held at Lawrence Berkeley National Laboratory (LBNL) on 27-29 May 2015. The goals of NDNCA were compile nuclear data needs across a wide spectrum of applied nuclear scie...
arxiv.org/abs/2501.02143v1
Safety-critical driving data is crucial for developing safe and trustworthy self-driving algorithms. Due to the scarcity of safety-critical data in naturalistic datasets, current approaches primarily utilize simulated or artificially generated images...
github.com/naturalistic-data-analysis/naturalistic_data_analysis
A jupyter book for the OHBM educational workshop on analyzing naturalistic data. (⭐ 46)
arxiv.org/abs/1710.10655v1
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are...