arxiv.org/abs/2505.07606v1
This article addresses the disconnect between the individual policy documents of Mastodon instances--many of which explicitly prohibit data collection for research purposes--and the actual data handling practices observed in academic research involvi...
arxiv.org/abs/1406.3766v1
We propose and preliminarily implement a data-mining based platform to assist experts to inspect the increasing amount of spectra with low signal to noise ratio (SNR) generated by large sky surveys. The platform includes three layers: data-mining lay...
github.com/MayumyCH/data-scientist-with-python-datacamp
Anotaciones del career "Data Scientist with Python" de Datacamp ?, gracias a la beca de DATASCIENCIEFEM?. #Data_Challenge_365_Fem ?? (⭐ 46)
arxiv.org/abs/1703.06450v1
Sharing scientific data, with the objective of making it fully discoverable, accessible, assessable, intelligible, usable, and interoperable, requires work at the disciplinary level to define in particular how the data should be formatted and describ...
arxiv.org/abs/2411.02247v1
The worlds of Data Science (including big and/or federated data, machine learning, etc) and Astrophysics started merging almost two decades ago. For instance, around 2005, international initiatives such as the Virtual Observatory framework rose to st...
arxiv.org/abs/2101.05638v1
La Serena School for Data Science is a multidisciplinary program with six editions so far and a constant format: during 10-14 days, a group of $\sim$30 students (15 from the US, 15 from Chile and 1-3 from Caribbean countries) and $\sim$9 faculty gath...
arxiv.org/abs/2412.11704v4
Vocabulary expansion (VE) is the de-facto approach to language adaptation of large language models (LLMs) by adding new tokens and continuing pre-training on target data. While this is effective for base models trained on unlabeled data, it poses cha...
arxiv.org/abs/2310.15799v1
We present DALE, a novel and effective generative Data Augmentation framework for low-resource LEgal NLP. DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents - legal language, with its...
www.bing.com/ck/a?!&&p=139eb680ea09346f1f0fcd03da80f849ce0aad9116f15ce7b3457c6c5abab2d7JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=223d71f2-c05e-678b-1fd5-66e7c1036609&u=a1aHR0cHM6Ly9vaWxwcmljZS5jb20vcmlnLWNvdW50&ntb=1
Sep 11, 2020 · Find U.S. and Canadian rig count and drilling data, and use Oilprice.com's graphing tools to compare oil prices, frac spread, production and drilling data per basin or province.
arxiv.org/abs/2407.13467v1
Six years after the entry into force of the GDPR, European companies and organizations still have difficulties complying with it: the amount of fines issued by the European data protection authorities is continuously increasing. Personal data transfe...
arxiv.org/abs/2407.02437v2
The widespread use of publicly available datasets for training machine learning models raises significant concerns about data misuse. Availability attacks have emerged as a means for data owners to safeguard their data by designing imperceptible pert...
www.bing.com/ck/a?!&&p=2db06845938153e2af11d6ec31e6f2078bb6ce89959ff207e2edfe4cc45452dfJmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=3a5fdec4-0e78-6f48-0d24-c9d10fd96eb1&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
github.com/stevespringett/nist-data-mirror
A simple Java command-line utility to mirror the CVE JSON data from NIST. (⭐ 214)
github.com/ashish-3916/Coding-Ninjas-Data-Structures
This repo contains solutions to problem of data structures in c++ (⭐ 209)
arxiv.org/abs/1704.01802v1
As part of Smart Cities initiatives, national, regional and local governments all over the globe are under the mandate of being more open regarding how they share their data. Under this mandate, many of these governments are publishing data under the...
en.wikipedia.org/wiki/Factor_analysis_of_mixed_data
mixed data (FAMD, in the French original: AFDM or Analyse Factorielle de Données Mixtes), is the factorial method devoted to data tables in which a group
www.bing.com/ck/a?!&&p=f2fff4aa2649cf4764279bf8c152d1c8967813712838d9e8dbcf42bb03726c56JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=304152eb-5841-6874-0826-45fe594a698f&u=a1aHR0cHM6Ly93d3cuYmVsbW9udGZvcnVtLm9yZy93cC1jb250ZW50L3VwbG9hZHMvMjAxOS8xMC8yLTctR3JhbnQtQXRsYW50T1MtRU1TTy1DT09QLnBkZg&ntb=1
Oct 2, 2019 · Encourage full meta data delivery with all data sets and to establish and promote the use standard descriptors to allow best data harvesting Include to the EOVs discussion issues such as …
www.bing.com/ck/a?!&&p=fc85f8cd3259839ed671a83d9ecb78759f32218e958838d1d80ec416b48f7553JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=0ba79ae2-0f30-6409-16fb-8df70e6565df&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
arxiv.org/abs/2303.10158v3
Artificial Intelligence (AI) is making a profound impact in almost every domain. A vital enabler of its great success is the availability of abundant and high-quality data for building machine learning models. Recently, the role of data in AI has bee...
arxiv.org/abs/2112.03837v1
Data scarcity and noise are important issues in industrial applications of machine learning. However, it is often challenging to devise a scalable and generalized approach to address the fundamental distributional and semantic properties of dataset w...