Cross Validated
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
Data Analytics Skills for Your Future Workplace Group Project Handbook Overview (⭐ 0)
As generation Z's big data is flooding the Internet through social nets, neural network based data processing is turning an important cornerstone, showing significant potential for fast extraction of data patterns. Online course delivery and associat...
structured data between applications, which was not its primary design goal. However, XML data binding systems allow applications to access XML data directly
The release of NBA player tracking data greatly enhances the granularity and dimensionality of basketball statistics used to evaluate and compare player performance. However, the high dimensionality of this new data source can be troublesome as it de...
A curated list of awesome posts, videos, and articles on leading a data team (small and large) (⭐ 549)
In this study, regional (cities, towns and villages) data and tweet data are obtained from Twitter, and extract information of purchase information (Where and what bought) from the tweet data by morphological analysis and rule-based dependency analys...
We study the following fundamental data-driven pricing problem. How can/should a decision-maker price its product based on data at a single historical price? How valuable is such data? We consider a decision-maker who optimizes over (potentially rand...
Spurred by developments such as cloud computing, there are increasing efforts for outsourcing of data management. A company (data owner) who lacks expertise and comptational resources can outsource his data to a third-party service provider (server),...
Jan 26, 2017 · I have following data and code to round selected columns of this data.table: mydf = structure (list (vnum1 = c (0.590165705411504, -1.39939534199836, 0.720226053660755, …
This paper describes a machine learning approach for annotating and analyzing data curation work logs at ICPSR, a large social sciences data archive. The systems we studied track curation work and coordinate team decision-making at ICPSR. Repository...
Data annotation remains the sine qua non of machine learning and AI. Recent empirical work on data annotation has begun to highlight the importance of rater diversity for fairness, model performance, and new lines of research have begun to examine th...
This paper introduces the new data-dependent multiplier bootstrap for non-parametric analysis of survival data, possibly subject to competing risks. The new resampling procedure includes both the general wild bootstrap and the weird bootstrap as spec...
Most scholars, politicians, and activists are following individualistic theories of privacy and data protection. In contrast, some of the pioneers of the data protection legislation in Germany like Adalbert Podlech, Paul J. Müller, and Ulrich Damman...
Codes for case studies for the Bekes-Kezdi Data Analysis textbook (⭐ 232)
Deep learning and other big data technologies have over time become very powerful and accurate. There are algorithms and models developed that have near human accuracy in their task. In health care, the amount of data available is massive and hence,...
We present an instrumenting compiler for enforcing data confidentiality in low-level applications (e.g. those written in C) in the presence of an active adversary. In our approach, the programmer marks secret data by writing lightweight annotations o...
A list of useful resources to learn Data Engineering from scratch (⭐ 3960)
The data found within the Pokémon TCG API (⭐ 741)
We present a catalog of 37,842 quasars in the SDSS Data Release 7, which have counterparts within 6" in the WISE Preliminary Data Release. The overall WISE detection rate of the SDSS quasars is 86.7%, and it decreases to less than 50.0% when the quas...