arxiv.org/abs/2507.23115v1
Previous work on data privacy in federated learning systems focuses on privacy-preserving operations for data from users who have agreed to share their data for training. However, modern data privacy agreements also empower users to use the system wh...
arxiv.org/abs/2404.15859v1
In light of the GDPR, data controllers (DC) need to allow data subjects (DS) to exercise certain data subject rights. A key requirement here is that DCs can reliably authenticate a DS. Due to a lack of clear technical specifications, this has been re...
github.com/andresvourakis/data-scientist-handbook
Curated Data Science resources (Free & Paid) to help aspiring and experienced data scientists learn, grow, and advance their careers. (⭐ 1400)
github.com/andrewgbruce/statistics-for-data-scientists
Code and data associated with the book "Statistics for Data Scientists: 50 Essential Concepts" (⭐ 1212)
github.com/YouGov-Data/covid-19-tracker
This is the data repository for the Imperial College London YouGov Covid 19 Behaviour Tracker Data Hub. (⭐ 92)
arxiv.org/abs/2308.09670v2
We use the computer algebra system GAP to classify modular data up to rank 12. This extends the previously obtained classification of modular data up to rank 6. Our classification includes all the modular data from modular tensor categories up to ran...
github.com/AleksiKnuutila/uk-data-controllers
UK Data Controllers, as Data Package, from UK information Commissioner's Office (ICO) (⭐ 0)
github.com/LuisM78/Appliances-energy-prediction-data
Data sets and scripts for the publication in Energy and Buildings Data driven prediction models of energy use of appliances in a low-energy house. Luis M. Candanedo, Véronique Feldheim, Dominique Deramaix. Energy and Buildings, Volume 140, 1 April 2017, Pages…
arxiv.org/abs/2309.12168v1
Data transformation is an essential step in data science. While experts primarily use programming to transform their data, there is an increasing need to support non-programmers with user interface-based tools. With the rapid development in interacti...
arxiv.org/abs/1711.04367v1
Analyzing patterns in data streams generated by network traffic, sensor networks, or satellite feeds is a challenge for systems in which the available storage is limited. In addition, real data is noisy, which makes designing data stream algorithms e...
arxiv.org/abs/0804.4071v1
Knowledge could be gained from experts, specialists in the area of interest, or it can be gained by induction from sets of data. Automatic induction of knowledge from data sets, usually stored in large databases, is called data mining. Data mining...
arxiv.org/abs/2503.17428v1
Data mining is not an invasion of privacy because access to data is only by machines, not by people: this is the argument that is investigated here. The current importance of this problem is developed in a case study of data mining in the USA for cou...
en.wikipedia.org/wiki/Missing_data
In statistics, missing data, or missing values, occur when no data value is stored for the variable in an observation. Missing data are a common occurrence
arxiv.org/abs/2104.03446v2
Massive data bring the big challenges of memory and computation for analysis. These challenges can be tackled by taking subsamples from the full data as a surrogate. For functional data, it is common to collect multiple measurements over their domain...
arxiv.org/abs/2009.06155v1
Over the last decade growing amounts of government data have been made available in an attempt to increase transparency and civic participation, but it is unclear if this data serves non-expert communities due to gaps in access and the technical know...
github.com/DataTalksClub/data-engineering-zoomcamp
Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here ?? (⭐ 38932)
github.com/firmai/data-science-career
Career Resources for Data Science, Machine Learning, Big Data and Business Analytics Career Repository (⭐ 1001)
www.bing.com/ck/a?!&&p=369467a9738d5949328dd02b9e8fbb378744a3c310ac28a7b8e2ee5495470a8fJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=0c2c6c99-301f-6424-03ca-7b8831f26504&u=a1aHR0cHM6Ly93d3cuYmVsbW9udGZvcnVtLm9yZy93cC1jb250ZW50L3VwbG9hZHMvMjAxOS8xMC9DUkFfRGF0YV9EaWdpdGFsX091dHB1dHNfTWFuYWdlbWVudF9WMi5wZGY&ntb=1
A full Data and Digital Outputs Management Plan for an awarded Belmont Forum project is a living, actively updated document that describes the data management life cycle for the data …
en.wikipedia.org/wiki/Data_scrubbing
then corrects detected errors using redundant data in the form of different checksums or copies of data. Data scrubbing reduces the likelihood that single
en.wikipedia.org/wiki/Data_science
knowledge to summarize data. Data science is an interdisciplinary field focused on extracting knowledge from typically large data sets and applying the knowledge