arxiv.org/abs/2502.14854v2
LLM developers are increasingly reliant on synthetic data, but generating high-quality data for complex long-context reasoning tasks remains challenging. We introduce CLIPPER, a compression-based approach for generating synthetic data tailored to nar...
arxiv.org/abs/2409.03741v1
Machine learning has revolutionized numerous domains, playing a crucial role in driving advancements and enabling data-centric processes. The significance of data in training models and shaping their performance cannot be overstated. Recent research...
arxiv.org/abs/1611.00481v2
In the era of big data, it is common to have data with multiple modalities or coming from multiple sources, known as "multi-view data". Multi-view clustering provides a natural way to generate clusters from such data. Since different views share some...
www.bing.com/ck/a?!&&p=3ba2a2c93fe88aba6e702908d78c01b178014584be60d2c6f4110ffc74479e53JmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=2e84aa39-ee0f-6170-0fcc-bd2bef3a605c&u=a1aHR0cHM6Ly9lbmdsaXNoLnN0YWNrZXhjaGFuZ2UuY29tL3F1ZXN0aW9ucy8xNjEwODEvd2hhdC1pcy1mcmVlLWZvcm0tZGF0YS1lbnRyeQ&ntb=1
If you are storing documents, however, you should choose either the mediumtext or longtext type. Could you please tell me what free-form data entry is? I know what data entry is per se - when data is fed …
arxiv.org/abs/1601.03115v1
Big Data can mean different things to different people. The scale and challenges of Big Data are often described using three attributes, namely Volume, Velocity and Variety (3Vs), which only reflect some of the aspects of data. In this chapter we rev...
arxiv.org/abs/2105.03161v1
In the Open Data Portal Germany (OPAL) project, a pipeline of the following data refinement steps has been developed: requirements analysis, data acquisition, analysis, conversion, integration and selection. 800,000 datasets in DCAT format have been...
arxiv.org/abs/1811.01429v2
Multivariate functional data are becoming ubiquitous with advances in modern technology and are substantially more complex than univariate functional data. We propose and study a novel model for multivariate functional data where the component proces...
arxiv.org/abs/2509.10165v2
Companies are looking to data anonymization research $\unicode{x2013}$ including differential private and synthetic data methods $\unicode{x2013}$ for simple and straightforward compliance solutions. But data anonymization has not taken off in practi...
arxiv.org/abs/1302.4133v1
NVD is one of the most popular databases used by researchers to conduct empirical research on data sets of vulnerabilities. Our recent analysis on Chrome vulnerability data reported by NVD has revealed an abnormally phenomenon in the data where almos...
github.com/ift-gftc/SeafoodTrackathon
A hackathon using real data sets taken from industry supply chains to develop solutions which address tools and applications to capture data on fishing vessels, conceptualize feasible data sharing options and develop creative ways to define data formats that s…
arxiv.org/abs/2308.16109v1
Accessibility of research data is critical for advances in many research fields, but textual data often cannot be shared due to the personal and sensitive information which it contains, e.g names or political opinions. General Data Protection Regulat...
arxiv.org/abs/1112.1668v1
Electronic health records (EHR's) are only a first step in capturing and utilizing health-related data - the problem is turning that data into useful information. Models produced via data mining and predictive analysis profile inherited risks and env...
arxiv.org/abs/1904.08932v1
Despite the benefits of school management information systems (SMIS), the concept of data-driven school culture failed to materialize for many educational institutions. Challenges posed by the quality of data in the big data era have prevented many s...
github.com/owid/covid-19-data
Data on COVID-19 (coronavirus) cases, deaths, hospitalizations, tests • All countries • Updated daily by Our World in Data (⭐ 5658)
arxiv.org/abs/2005.12532v1
The CMS experiment recorded 177.75 /fb of proton-proton collision data during the RUN-1 and RUN-2 data taking period. Successful data taking at increasing instantaneous luminosities with the evolving detector configuration was a big achievement of th...
arxiv.org/abs/1609.00031v1
Partially observed cured data occur in the analysis of spontaneous abortion (SAB) in observational studies in pregnancy. In contrast to the traditional cured data, such data has an observable `cured' portion as women who do not abort spontaneously. T...
arxiv.org/abs/1609.01810v1
To deal with many pedestrian data, automatic data collection is needed. This paper describes how to automate the microscopic pedestrian flow data collection from video files. The study is restricted only to pedestrians without considering vehicular -...
en.wikipedia.org/wiki/Data_haven
A data haven, like a corporate haven or tax haven, is a refuge for uninterrupted or unregulated data. Data havens are locations with legal environments
arxiv.org/abs/2601.21706v2
Smart meter data is the foundation for planning and operating the distribution network. Unfortunately, such data are not always available due to privacy regulations. Meanwhile, the collected data may be corrupted due to sensor or transmission failure...
arxiv.org/abs/2003.13822v1
While digital trace data from sources like search engines hold enormous potential for tracking and understanding human behavior, these streams of data lack information about the actual experiences of those individuals generating the data. Moreover, m...