arxiv.org/abs/2307.00161v1
Generative modeling has been used frequently in synthetic data generation. Fairness and privacy are two big concerns for synthetic data. Although Recent GAN [\cite{goodfellow2014generative}] based methods show good results in preserving privacy, the...
arxiv.org/abs/1907.06183v1
Today, the operating TAIGA (Tunka Advanced Instrument for cosmic rays and Gamma Astronomy) experiment continuously produces and accumulates a large volume of raw astroparticle data. To be available for the scientific community these data should be we...
arxiv.org/abs/2511.17694v1
This article describes the use of metadata and standards in the Social Impact Data Commons to expose official statisticians to an innovative project built on actionable and evaluable metadata, which produces a FAIR data system. We begin by introducin...
github.com/open-metadata/OpenMetadata
OpenMetadata is a unified metadata platform for data discovery, data observability, and data governance powered by a central metadata repository, in-depth column level lineage, and seamless team collaboration. (⭐ 8859)
www.bing.com/ck/a?!&&p=7c5b389c48bc8565a8ff76b04d0d8309f3c6f91f5271f7d02bfd7a938def0dbcJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=208fbb1a-248f-66b9-06e8-ac0b253c6779&u=a1aHR0cHM6Ly93d3cuaWJtLmNvbS90aGluay90b3BpY3MvbWV0YWRhdGE&ntb=1
What is metadata? Metadata is information—such as author, creation date or file size—that describes a data point or data set. Metadata can improve a data system’s functions and make it easier to search …
en.wikipedia.org/wiki/Metadata
about subject descriptions of data and token codes for the data. We also have statements in a meta language describing the data relationships and transformations
arxiv.org/abs/2003.00899v2
Most datasets of interest to the analytics industry are impacted by various forms of human bias. The outcomes of Data Analytics [DA] or Machine Learning [ML] on such data are therefore prone to replicating the bias. As a result, a large number of bia...
arxiv.org/abs/2012.00697v3
The new generation of cloud data warehouses (CDWs) brings large amounts of data and compute power closer to users in enterprises. The ability to directly access the warehouse data, interactively analyze and explore it at scale can empower users to im...
arxiv.org/abs/2404.00776v2
We present PyTorch Frame, a PyTorch-based framework for deep learning over multi-modal tabular data. PyTorch Frame makes tabular deep learning easy by providing a PyTorch-based data structure to handle complex tabular data, introducing a model abstra...
arxiv.org/abs/2009.02181v2
Wireless data aggregation (WDA), referring to aggregating data distributed at devices (e.g., sensors and smartphone), is a common operation in 5G-and-beyond machine-type communications to support Internet-of-Things (IoT), which lays the foundation fo...
arxiv.org/abs/1805.08359v2
In the past few years, we have envisioned an increasing number of businesses start driving by big data analytics, such as Amazon recommendations and Google Advertisements. At the back-end side, the businesses are powered by big data processing platfo...
arxiv.org/abs/1810.06765v1
Use of medical data, also known as electronic health records, in research helps develop and advance medical science. However, protecting patient confidentiality and identity while using medical data for analysis is crucial. Medical data can be in the...
arxiv.org/abs/2508.18892v1
These recommendations are the result of reflections by scientists and experts who are, or have been, involved in the preservation of high-energy physics data. The work has been done under the umbrella of the Data Lifecycle panel of the International...
github.com/RfastOfficial/Rfast2
A collection of Rfast2 functions for data analysis. Note 1: The vast majority of the functions accept matrices only, not data.frames. Note 2: Do not have matrices or vectors with have missing data (i.e NAs). We do no check about them and C++ internally transfo…
github.com/jeremiak/us-senate-financial-disclosure-data
US Senate financial reports and structured data (⭐ 12)
arxiv.org/abs/1808.02528v3
Randomized evaluations of educational technology produce log data as a bi-product: highly granular data student and teacher usage. These datasets could shed light on causal mechanisms, effect heterogeneity, or optimal use. However, there are methodol...
arxiv.org/abs/0912.0255v1
Data from high-energy physics (HEP) experiments are collected with significant financial and human effort and are mostly unique. At the same time, HEP has no coherent strategy for data preservation and re-use. An inter-experimental Study Group on H...
github.com/nychealth/coronavirus-data
This repository contains data on Coronavirus Disease 2019 (COVID-19) in New York City (NYC), from the NYC Department of Health and Mental Hygiene. (⭐ 954)
github.com/jamiebuilds/itsy-bitsy-data-structures
:european_castle: All the things you didn't know you wanted to know about data structures (⭐ 8585)
github.com/qusaybtoush/Texas-Wind---Turbine---Accuracy-99-
Texas Wind - Turbine About Dataset Problem Statement: The intermittent nature and low control over the wind conditions bring up the same problem to every grid operator in their successful integration to satisfy current demand. In combination with having to pre…