arxiv.org/abs/2102.04072v3
We propose Blue Noise Plots, two-dimensional dot plots that depict data points of univariate data sets. While often one-dimensional strip plots are used to depict such data, one of their main problems is visual clutter which results from overlap. To...
github.com/democrats/data
Datasets that list various boolean, categorical, and date data for elections and elected officials. (⭐ 65)
github.com/harvard-lil/data-vault
Tools for LIL's data preservation project (⭐ 126)
arxiv.org/abs/2509.01774v2
Clustered and longitudinal data are pervasive in scientific studies, from prenatal health programs to clinical trials and public health surveillance. Such data often involve non-Gaussian responses--including binary, categorical, and count outcomes--t...
arxiv.org/abs/2502.05191v1
Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific cli...
arxiv.org/abs/2207.06253v1
In modern data science, it is common that large-scale data are stored and processed parallelly across a great number of locations. For reasons including confidentiality concerns, only limited data information from each parallel center is eligible to...
www.bing.com/ck/a?!&&p=e4a11bad3150fb6135ace250617c35b55aea178eadf6de9c92fa1476a421be7eJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=39abaab2-d144-6dba-3dd1-bda3d0496c8b&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
www.bing.com/ck/a?!&&p=8edf93cda8c35faad48d84afebd556d09a6f1077e339f91d5ec3396703f42c6cJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=115bb4ff-8149-61f9-0835-a3ee805560e9&u=a1aHR0cHM6Ly9lbi53aWtpcGVkaWEub3JnL3dpa2kvR3JhcGhpY3M&ntb=1
A graph or chart is a graphic that represents tabular or numeric data. Charts are often used to make it easier to understand large quantities of data and the relationships between different parts of the data.
arxiv.org/abs/1508.02076v2
We present the GAMA Panchromatic Data Release (PDR) constituting over 230deg$^2$ of imaging with photometry in 21 bands extending from the far-UV to the far-IR. These data complement our spectroscopic campaign of over 300k galaxies, and are compiled...
arxiv.org/abs/1901.03723v1
An increasing number of cybersecurity incidents prompts organizations to explore alternative security solutions, such as threat intelligence programs. For such programs to succeed, data needs to be collected, validated, and recorded in relevant datas...
arxiv.org/abs/1502.02057v1
We examine Illinois educational data from standardized exams and analyze primary factors affecting the achievement of public school students. We focus on the simplest possible models: representation of data through visualizations and regressions on s...
arxiv.org/abs/1907.10173v1
We investigate trend identification in the LML and MAN atmospheric ammonia data. The signals are mixed in the LML data, with just as many positive, negative, and no trends found. The start date for trend identification is crucial, with the trends cla...
arxiv.org/abs/2602.06772v1
Data augmentation can mitigate limited training data in machine-learning automated scoring engines for constructed response items. This study seeks to determine how well three approaches to large language model prompting produce essays that preserve...
github.com/spring-tips/data-oriented-programming-in-java-21
Hi, Spring fans! In this installment we look at one of my favorite paradigms in Java 21 and later: Data Oriented Programming (⭐ 11)
arxiv.org/abs/2412.05055v4
Increasing numbers of athletes and sports teams use data collection technologies to improve athletic development and athlete health with the goal of improving competitive performance. Personal data privacy is managed but it is not always a priority f...
github.com/garmin-data/garmdown
Download Garmin Connect Data (⭐ 17)
arxiv.org/abs/1809.09081v1
The generative learning phase of Autoencoder (AE) and its successor Denosing Autoencoder (DAE) enhances the flexibility of data stream method in exploiting unlabelled samples. Nonetheless, the feasibility of DAE for data stream analytic deserves in-d...
en.wikipedia.org/wiki/Data_monitoring_committee
A data monitoring committee (DMC) – sometimes called a data and safety monitoring board (DSMB) – is an independent group of experts who monitor patient
arxiv.org/abs/1302.4381v3
We extend the theory of d-separation to cases in which data instances are not independent and identically distributed. We show that applying the rules of d-separation directly to the structure of probabilistic models of relational data inaccurately i...
arxiv.org/abs/2102.13125v1
Current Cloud solutions for Edge Computing are inefficient for data-centric applications, as they focus on the IaaS/PaaS level and they miss the data modeling and operations perspective. Consequently, Edge Computing opportunities are lost due to cumb...