arxiv.org/abs/2411.02247v1
The worlds of Data Science (including big and/or federated data, machine learning, etc) and Astrophysics started merging almost two decades ago. For instance, around 2005, international initiatives such as the Virtual Observatory framework rose to st...
arxiv.org/abs/2101.05638v1
La Serena School for Data Science is a multidisciplinary program with six editions so far and a constant format: during 10-14 days, a group of $\sim$30 students (15 from the US, 15 from Chile and 1-3 from Caribbean countries) and $\sim$9 faculty gath...
arxiv.org/abs/2412.11704v4
Vocabulary expansion (VE) is the de-facto approach to language adaptation of large language models (LLMs) by adding new tokens and continuing pre-training on target data. While this is effective for base models trained on unlabeled data, it poses cha...
arxiv.org/abs/2310.15799v1
We present DALE, a novel and effective generative Data Augmentation framework for low-resource LEgal NLP. DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents - legal language, with its...
www.bing.com/ck/a?!&&p=139eb680ea09346f1f0fcd03da80f849ce0aad9116f15ce7b3457c6c5abab2d7JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=223d71f2-c05e-678b-1fd5-66e7c1036609&u=a1aHR0cHM6Ly9vaWxwcmljZS5jb20vcmlnLWNvdW50&ntb=1
Sep 11, 2020 · Find U.S. and Canadian rig count and drilling data, and use Oilprice.com's graphing tools to compare oil prices, frac spread, production and drilling data per basin or province.
arxiv.org/abs/2407.13467v1
Six years after the entry into force of the GDPR, European companies and organizations still have difficulties complying with it: the amount of fines issued by the European data protection authorities is continuously increasing. Personal data transfe...
arxiv.org/abs/2407.02437v2
The widespread use of publicly available datasets for training machine learning models raises significant concerns about data misuse. Availability attacks have emerged as a means for data owners to safeguard their data by designing imperceptible pert...
www.bing.com/ck/a?!&&p=2db06845938153e2af11d6ec31e6f2078bb6ce89959ff207e2edfe4cc45452dfJmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=3a5fdec4-0e78-6f48-0d24-c9d10fd96eb1&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
github.com/stevespringett/nist-data-mirror
A simple Java command-line utility to mirror the CVE JSON data from NIST. (⭐ 214)
github.com/ashish-3916/Coding-Ninjas-Data-Structures
This repo contains solutions to problem of data structures in c++ (⭐ 209)
arxiv.org/abs/1704.01802v1
As part of Smart Cities initiatives, national, regional and local governments all over the globe are under the mandate of being more open regarding how they share their data. Under this mandate, many of these governments are publishing data under the...
en.wikipedia.org/wiki/Factor_analysis_of_mixed_data
mixed data (FAMD, in the French original: AFDM or Analyse Factorielle de Données Mixtes), is the factorial method devoted to data tables in which a group
www.bing.com/ck/a?!&&p=f2fff4aa2649cf4764279bf8c152d1c8967813712838d9e8dbcf42bb03726c56JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=304152eb-5841-6874-0826-45fe594a698f&u=a1aHR0cHM6Ly93d3cuYmVsbW9udGZvcnVtLm9yZy93cC1jb250ZW50L3VwbG9hZHMvMjAxOS8xMC8yLTctR3JhbnQtQXRsYW50T1MtRU1TTy1DT09QLnBkZg&ntb=1
Oct 2, 2019 · Encourage full meta data delivery with all data sets and to establish and promote the use standard descriptors to allow best data harvesting Include to the EOVs discussion issues such as …
www.bing.com/ck/a?!&&p=fc85f8cd3259839ed671a83d9ecb78759f32218e958838d1d80ec416b48f7553JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=0ba79ae2-0f30-6409-16fb-8df70e6565df&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
arxiv.org/abs/2303.10158v3
Artificial Intelligence (AI) is making a profound impact in almost every domain. A vital enabler of its great success is the availability of abundant and high-quality data for building machine learning models. Recently, the role of data in AI has bee...
arxiv.org/abs/2112.03837v1
Data scarcity and noise are important issues in industrial applications of machine learning. However, it is often challenging to devise a scalable and generalized approach to address the fundamental distributional and semantic properties of dataset w...
www.bing.com/ck/a?!&&p=39f1f96eb8dddbfb836c88eaad60326c722f0371f30fc2f0894dbdbdb7b01063JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=3e2dcbb3-36e6-67e0-110d-dca637e7662a&u=a1aHR0cHM6Ly9vaWxwcmljZS5jb20vcmlnLWNvdW50&ntb=1
Sep 11, 2020 · Find U.S. and Canadian rig count and drilling data, and use Oilprice.com's graphing tools to compare oil prices, frac spread, production and drilling data per basin or province.
arxiv.org/abs/2404.04299v1
Summary: The vast generation of genetic data poses a significant challenge in efficiently uncovering valuable knowledge. Introducing GENEVIC, an AI-driven chat framework that tackles this challenge by bridging the gap between genetic data generation...
arxiv.org/abs/2005.00388v1
We present a new challenging stance detection dataset, called Will-They-Won't-They (WT-WT), which contains 51,284 tweets in English, making it by far the largest available dataset of the type. All the annotations are carried out by experts; therefore...
arxiv.org/abs/2405.18153v3
Machine Listening focuses on developing technologies to extract relevant information from audio signals. A critical aspect of these projects is the acquisition and labeling of contextualized data, which is inherently complex and requires specific res...