arxiv.org/abs/0912.0255v1
Data from high-energy physics (HEP) experiments are collected with significant financial and human effort and are mostly unique. At the same time, HEP has no coherent strategy for data preservation and re-use. An inter-experimental Study Group on H...
arxiv.org/abs/2211.09098v1
The vehicle recognition area, including vehicle make-model recognition (VMMR), re-id, tracking, and parts-detection, has made significant progress in recent years, driven by several large-scale datasets for each task. These datasets are often non-ove...
github.com/Voce-lin/Robin-time-series-dataset
Time series anomaly detection dataset from an anonymous commercial bank (⭐ 0)
github.com/nychealth/coronavirus-data
This repository contains data on Coronavirus Disease 2019 (COVID-19) in New York City (NYC), from the NYC Department of Health and Mental Hygiene. (⭐ 954)
github.com/jamiebuilds/itsy-bitsy-data-structures
:european_castle: All the things you didn't know you wanted to know about data structures (⭐ 8585)
github.com/qusaybtoush/Texas-Wind---Turbine---Accuracy-99-
Texas Wind - Turbine About Dataset Problem Statement: The intermittent nature and low control over the wind conditions bring up the same problem to every grid operator in their successful integration to satisfy current demand. In combination with having to pre…
arxiv.org/abs/1901.11040v1
Benchmark data sets are of vital importance in machine learning research, as indicated by the number of repositories that exist to make them publicly available. Although many of these are usable in the stream mining context as well, it is less obviou...
arxiv.org/abs/1801.06258v1
This paper addresses the Data-Diff problem: given a dataset and a subsequent version of the dataset, find the shortest sequence of operations that transforms the dataset to the subsequent version, under a restricted family of operations. We consider...
arxiv.org/abs/1909.10631v3
Most historical National Football League (NFL) analysis, both mainstream and academic, has relied on public, play-level data to generate team and player comparisons. Given the number of oft omitted variables that impact on-field results, such as play...
github.com/Skandy01/SQL-Walmart-Data-Analysis-Project
Spearheaded an exhaustive analysis of Walmart sales data sourced from the Kaggle Walmart Sales Forecasting Competition.Identified optimal branch locations, resulting in a 5% reduction in operating costs. Analyzed payment methods, with "Credit Card" emerging as…
arxiv.org/abs/1609.08766v1
Data is a dominant force during the decision-making process. It can help determine which roads to expand and the optimal location for a grocery store. Data can also be used to influence which schools to open or to shutter and "appropriate" city servi...
arxiv.org/abs/1708.07759v1
Earlier attempts to investigate the changes of the role of friendship in different life stages have failed due to lack of data. We close this gap by using a large data set of mobile phone calls from a European country in 2007, to study how the people...
www.reddit.com/r/devsarg/comments/1q8k4fr/si_compro_un_curso_de_data_analyst_en_udemy/
Me gusta el data analyst y estuve estudiando inglés un tiempo para poder meterme como se debe en IT como data scientist pero ando necesitando chamba así que creo que no debería esperar más y empez...
arxiv.org/abs/1805.11457v2
Here is a database of quasicrystal cells computed by the deBruijn Grand Dual Method. The database is in a form that can be converted and read by a variety of geometry programs. Proof of the accuracy of the computations is given by the consistency of...
arxiv.org/abs/2406.02623v2
Background: Limited universally-adopted data standards in veterinary medicine hinder data interoperability and therefore integration and comparison; this ultimately impedes the application of existing information-based tools to support advancement in...
arxiv.org/abs/1709.01989v1
Data science and machine learning are the key technologies when it comes to the processes and products with automatic learning and optimization to be used in the automotive industry of the future. This article defines the terms "data science" (also r...
www.reddit.com/r/Epstein/comments/1qu77e4/massively_breaking_news_dataset_9_of_epstein/
The file in question was taken down by the DOJ but it was originally found here: [https://www.justice.gov/epstein/files/DataSet%209/EFTA01250883.pdf](https://www.justice.gov/epstein/files/DataSet%209/...
www.reddit.com/r/sidestreetbets/comments/1qu7429/massively_breaking_news_dataset_9_of_epstein/
The file in question was taken down by the DOJ but it was originally found here: [https://www.justice.gov/epstein/files/DataSet%209/EFTA01250883.pdf](https://www.justice.gov/epstein/files/DataSet%209/...
github.com/zhaohengyang/Generate-Parallel-Data-for-Sentence-Compression
Implement Overcoming the Lack of Parallel Data in Sentence Compression Katja Filippova and Yasemin Altun Google (⭐ 14)
arxiv.org/abs/2601.06057v2
The report highlights the role of Egyptian data workers in the global value chains of Artificial Intelligence (AI). These workers generate and annotate data for machine learning, check outputs, and they connect with overseas AI producers via internat...