arxiv.org/abs/1604.02533v1
This paper studies two design tasks faced by a geo-distributed cloud data market: which data to purchase (data purchasing) and where to place/replicate the data for delivery (data placement). We show that the joint problem of data purchasing and data...
arxiv.org/abs/2212.09849v6
Fine-tuning pre-trained language models has become the prevalent paradigm for building downstream NLP models. Oftentimes fine-tuned models are readily available but their training data is not, due to data privacy or intellectual property concerns. Th...
arxiv.org/abs/2410.10270v3
Discovering meaningful insights from a large dataset, known as Exploratory Data Analysis (EDA), is a challenging task that requires thorough exploration and analysis of the data. Automated Data Exploration (ADE) systems use goal-oriented methods with...
arxiv.org/abs/1905.07002v2
Large-scale clinical data is invaluable to driving many computational scientific advances today. However, understandable concerns regarding patient privacy hinder the open dissemination of such data and give rise to suboptimal siloed research. De-ide...
arxiv.org/abs/1407.7795v1
Many data stewards collect confidential data that include fine geography. When sharing these data with others, data stewards strive to disseminate data that are informative for a wide range of spatial and non-spatial analyses while simultaneously pro...
arxiv.org/abs/1710.08874v1
Data for good implies unfettered access to data. But data owners must be conservative about how, when, and why they share data or risk violating the trust of the people they aim to help, losing their funding, or breaking the law. Data sharing agreeme...
arxiv.org/abs/2602.23463v1
The use of synthetic data has emerged as an essential tool in Magnetic Resonance Spectroscopy (MRS) research and applications, providing advantages for optimization of acquisition, software validation, deep learning applications, and enhanced reprodu...
arxiv.org/abs/2311.17453v1
Synthetic data (SD) have garnered attention as a privacy enhancing technology. Unfortunately, there is no standard for quantifying their degree of privacy protection. In this paper, we discuss proposed quantification approaches. This contributes to t...
arxiv.org/abs/2206.06215v2
Gaia Data Release 3 provides novel flux-calibrated low-resolution spectrophotometry for about 220 million sources in the wavelength range 330nm - 1050nm (XP spectra). Synthetic photometry directly tied to a flux in physical units can be obtained from...
arxiv.org/abs/2601.14791v1
The scarcity of training data presents a fundamental challenge in applying deep learning to archaeological artifact classification, particularly for the rare types of Chinese porcelain. This study investigates whether synthetic images generated throu...
arxiv.org/abs/1411.6273v1
In many simulation studies involving networks there is the need to rely on a sample network to perform the simulation experiments. In many cases, real network data is not available due to privacy concerns. In that case we can recourse to synthetic da...
www.reddit.com/r/MachineLearning/comments/1ew12xp/p_anyclassifier_synthetic_data_generation_for/
I would like to share this with everyone, as I thought it would be a great resource for most ML engineer and software engineers. I created a synthetic data generation for text classification module. ...
arxiv.org/abs/2210.03529v2
Recent advances in synthesizing realistic faces have shown that synthetic training data can replace real data for various face-related computer vision tasks. A question arises: how important is realism? Is the pursuit of photorealism excessive? In th...
arxiv.org/abs/2601.12124v1
The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks,...
arxiv.org/abs/2401.01734v2
Assistive robots should be able to wash, fold or iron clothes. However, due to the variety, deformability and self-occlusions of clothes, creating robot systems for cloth manipulation is challenging. Synthetic data is a promising direction to improve...
arxiv.org/abs/2208.07337v1
This paper presents a summary of the Competition on Face Morphing Attack Detection Based on Privacy-aware Synthetic Training Data (SYN-MAD) held at the 2022 International Joint Conference on Biometrics (IJCB 2022). The competition attracted a total o...
arxiv.org/abs/2309.05950v5
Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However, many VLMs rely on proprietary data and are not open-source, which restricts the...
developer.android.com/guide/topics/large-screens/get-started-with-large-screens
Large screens expand your app development opportunities. The large screens of tablets, foldables, and ChromeOS devices showcase content, facilitate multitasking, and enable user interfaces not possible on small screens.
developer.android.com//guide/topics/large-screens/get-started-with-large-screens
Large screens expand your app development opportunities. The large screens of tablets, foldables, and ChromeOS devices showcase content, facilitate multitasking, and enable user interfaces not possible on small screens.
developer.android.com/guide/topics/large-screens/get-started-with-large-screens
Large screens expand your app development opportunities. The large screens of tablets, foldables, and ChromeOS devices showcase content, facilitate multitasking, and enable user interfaces not possible on small screens.