arxiv.org/abs/2102.04072v3
We propose Blue Noise Plots, two-dimensional dot plots that depict data points of univariate data sets. While often one-dimensional strip plots are used to depict such data, one of their main problems is visual clutter which results from overlap. To...
github.com/democrats/data
Datasets that list various boolean, categorical, and date data for elections and elected officials. (⭐ 65)
github.com/harvard-lil/data-vault
Tools for LIL's data preservation project (⭐ 126)
github.com/github/advisory-database
Security vulnerability database inclusive of CVEs and GitHub originated security advisories from the world of open source software. (⭐ 2175)
arxiv.org/abs/2509.01774v2
Clustered and longitudinal data are pervasive in scientific studies, from prenatal health programs to clinical trials and public health surveillance. Such data often involve non-Gaussian responses--including binary, categorical, and count outcomes--t...
arxiv.org/abs/1801.07779v1
This paper describes the WiLI-2018 benchmark dataset for monolingual written natural language identification. WiLI-2018 is a publicly available, free of charge dataset of short text extracts from Wikipedia. It contains 1000 paragraphs of 235 language...
arxiv.org/abs/2502.05191v1
Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific cli...
arxiv.org/abs/2207.06253v1
In modern data science, it is common that large-scale data are stored and processed parallelly across a great number of locations. For reasons including confidentiality concerns, only limited data information from each parallel center is eligible to...
arxiv.org/abs/2308.16900v4
We present WineSensed, a large multimodal wine dataset for studying the relations between visual perception, language, and flavor. The dataset encompasses 897k images of wine labels and 824k reviews of wines curated from the Vivino platform. It has o...
www.bing.com/ck/a?!&&p=e4a11bad3150fb6135ace250617c35b55aea178eadf6de9c92fa1476a421be7eJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=39abaab2-d144-6dba-3dd1-bda3d0496c8b&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
github.com/spMohanty/PlantVillage-Dataset
Dataset of diseased plant leaf images and corresponding labels (⭐ 823)
github.com/jishanshaikh4/chess-games-dataset
Simplified dataset (.txt) of 50000+ chess games played online by Magnus Carlsen (⭐ 6)
www.bing.com/ck/a?!&&p=8edf93cda8c35faad48d84afebd556d09a6f1077e339f91d5ec3396703f42c6cJmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=115bb4ff-8149-61f9-0835-a3ee805560e9&u=a1aHR0cHM6Ly9lbi53aWtpcGVkaWEub3JnL3dpa2kvR3JhcGhpY3M&ntb=1
A graph or chart is a graphic that represents tabular or numeric data. Charts are often used to make it easier to understand large quantities of data and the relationships between different parts of the data.
en.wikipedia.org/wiki/Active_database
In computing, an active database is a database that includes an event-driven architecture (often in the form of ECA rules) that can respond to conditions
arxiv.org/abs/cs/0003004v1
Since scripts were proposed in the 1970's as an inferencing mechanism for AI and natural language processing programs, there have been few attempts to build a database of scripts. This paper describes a database and lexicon of scripts that has been...
arxiv.org/abs/1508.02076v2
We present the GAMA Panchromatic Data Release (PDR) constituting over 230deg$^2$ of imaging with photometry in 21 bands extending from the far-UV to the far-IR. These data complement our spectroscopic campaign of over 300k galaxies, and are compiled...
arxiv.org/abs/1901.03723v1
An increasing number of cybersecurity incidents prompts organizations to explore alternative security solutions, such as threat intelligence programs. For such programs to succeed, data needs to be collected, validated, and recorded in relevant datas...
arxiv.org/abs/1502.02057v1
We examine Illinois educational data from standardized exams and analyze primary factors affecting the achievement of public school students. We focus on the simplest possible models: representation of data through visualizations and regressions on s...
arxiv.org/abs/1907.10173v1
We investigate trend identification in the LML and MAN atmospheric ammonia data. The signals are mixed in the LML data, with just as many positive, negative, and no trends found. The start date for trend identification is crucial, with the trends cla...
arxiv.org/abs/2602.06772v1
Data augmentation can mitigate limited training data in machine-learning automated scoring engines for constructed response items. This study seeks to determine how well three approaches to large language model prompting produce essays that preserve...