arxiv.org/abs/1902.11104v3
When data are organized in matrices or arrays of higher dimensions (tensors), classical regression methods first transform these data into vectors, therefore ignoring the underlying structure of the data and increasing the dimensionality of the probl...
github.com/datacamp/course-resources-ml-with-experts-budgets
Further student resources for DrivenData's 'Machine Learning with the Experts: School Budgets' DataCamp course. (⭐ 553)
arxiv.org/abs/2307.06089v1
User Experience (UX) professionals need to be able to analyze large amounts of usage data on their own to make evidence-based design decisions. However, the design process for In-Vehicle Information Systems (IVIS) lacks data-driven support and effect...
www.bing.com/ck/a?!&&p=e57512748af7cbba55cd13e889518f8f6a17f08ce94cca36f6cd97e31d4c44d6JmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=09f6e0dc-77fb-6493-25b7-f7cd76ab6598&u=a1aHR0cHM6Ly9zdGF0cy5zdGFja2V4Y2hhbmdlLmNvbS8&ntb=1
Q&A for people interested in statistics, machine learning, data analysis, data mining, and data visualization
arxiv.org/abs/2511.07286v1
We present Glioma C6, a new open dataset for instance segmentation of glioma C6 cells, designed as both a benchmark and a training resource for deep learning models. The dataset comprises 75 high-resolution phase-contrast microscopy images with over...
en.wikipedia.org/wiki/Run-length_encoding
Run-length encoding (RLE) is a form of lossless data compression in which runs of data (consecutive occurrences of the same data value) are stored as a
arxiv.org/abs/1406.2104v1
The Daya Bay Reactor Neutrino Experiment started running on September 23, 2011. The offline computing environment, consisting of 11 servers at Daya Bay, was built to process onsite data. With current computing ability, onsite data processing is runni...
github.com/Ademdayan/Dogal-Dil-isleme-Turkiye-haberlerinden-
Preprocessing and cleaning processes are performed to retrieve news from the data set and select specific words, then the data is tokenized, then the data is trained and the result is obtained. - Veri setinden haberleri alıp belirli kelimeleri almak için önişl…
arxiv.org/abs/2209.09724v1
The Advanced Data Protection Control (ADPC) is a technical specification - and a set of sociotechnical mechanisms surrounding it - that can change the current practice of Internet-based personal data protection and consenting by providing novel and s...
arxiv.org/abs/0711.2861v1
We propose a maximum entropy (ME) based approach to smooth noise not only in data but also to noise amplified by second order derivative calculation of the data especially for electroencephalography (EEG) studies. The approach includes two steps, a...
github.com/Apress/spring-cloud-data-flow
Source Code for 'Spring Cloud Data Flow' by Felipe Gutierrez (⭐ 12)
github.com/kojino/120-Data-Science-Interview-Questions
Answers to 120 commonly asked data science interview questions. (⭐ 3835)
en.wikipedia.org/wiki/Data_entry
Data entry is the process of digitizing data by entering it into a computer system for organization and management purposes. It is a person-based process
arxiv.org/abs/1905.13345v3
We study the use of power weighted shortest path distance functions for clustering high dimensional Euclidean data, under the assumption that the data is drawn from a collection of disjoint low dimensional manifolds. We argue, theoretically and exper...
arxiv.org/abs/0707.4643v4
We develop an active set algorithm for the maximum likelihood estimation of a log-concave density based on complete data. Building on this fast algorithm, we indidate an EM algorithm to treat arbitrarily censored or binned data....
github.com/TrainingByPackt/Data-Wrangling-with-Python
Simplify your ETL processes with these hands-on data sanitation tips, tricks, and best practices (⭐ 132)
arxiv.org/abs/2505.22065v1
This paper presents the AquaMonitor dataset, the first large computer vision dataset of aquatic invertebrates collected during routine environmental monitoring. While several large species identification datasets exist, they are rarely collected usin...
arxiv.org/abs/2507.20329v1
Handling missing data is a major challenge in model-based clustering, especially when the data exhibit skewness and heavy tails. We address this by extending the finite mixture of scale mixtures of multivariate skew-normal (FMSMSN) family to accommod...
arxiv.org/abs/2401.11012v1
We are in the process of creating a database of digital games from the DACH region. This article provides an insight into the context in which it was created and the underlying methodological considerations behind the games database. The database was...
arxiv.org/abs/2312.10093v1
Record linkage means linking data from multiple sources. This approach enables the answering of scientific questions that cannot be addressed using single data sources due to limited variables. The potential of linked data for health research is enor...