arxiv.org/abs/2305.15723v1
In this paper, we study the setting in which data owners train machine learning models collaboratively under a privacy notion called joint differential privacy [Kearns et al., 2018]. In this setting, the model trained for each data owner $j$ uses $j$...
arxiv.org/abs/2106.05468v2
Vertical Federated Learning (VFL) refers to the collaborative training of a model on a dataset where the features of the dataset are split among multiple data owners, while label information is owned by a single data owner. In this paper, we propose...
arxiv.org/abs/1810.06936v2
Data-driven algorithms have surpassed traditional techniques in almost every aspect in robotic vision problems. Such algorithms need vast amounts of quality data to be able to work properly after their training process. Gathering and annotating that...
arxiv.org/abs/2412.06288v2
The surging demand for AI has led to a rapid expansion of energy-intensive data centers, impacting the environment through escalating carbon emissions and water consumption. While significant attention has been paid to data centers' growing environme...
arxiv.org/abs/2002.07397v1
Open-domain retrieval-based dialogue systems require a considerable amount of training data to learn their parameters. However, in practice, the negative samples of training data are usually selected from an unannotated conversation data set at rando...
arxiv.org/abs/2105.14284v3
Management of invasive species and pathogens requires information about the traffic of potential vectors. Such information is often taken from vector traffic models fitted to survey data. Here, user-specific data collected via mobile apps offer new o...
arxiv.org/abs/1511.07772v1
The Workshop on Nuclear Data Needs and Capabilities for Applications (NDNCA) was held at Lawrence Berkeley National Laboratory (LBNL) on 27-29 May 2015. The goals of NDNCA were compile nuclear data needs across a wide spectrum of applied nuclear scie...
arxiv.org/abs/2501.02143v1
Safety-critical driving data is crucial for developing safe and trustworthy self-driving algorithms. Due to the scarcity of safety-critical data in naturalistic datasets, current approaches primarily utilize simulated or artificially generated images...
github.com/naturalistic-data-analysis/naturalistic_data_analysis
A jupyter book for the OHBM educational workshop on analyzing naturalistic data. (⭐ 46)
arxiv.org/abs/1710.10655v1
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are...
arxiv.org/abs/1809.02869v3
Machine teaching addresses the problem of finding the best training data that can guide a learning algorithm to a target model with minimal effort. In conventional settings, a teacher provides data that are consistent with the true data distribution....
arxiv.org/abs/1903.12525v1
Bloom Filter is a probabilistic membership data structure and it is excessively used data structure for membership query. Bloom Filter becomes the predominant data structure in approximate membership filtering. Bloom Filter extremely enhances the que...
arxiv.org/abs/1501.01941v3
Bloom filters are probabilistic data structures commonly used for approximate membership problems in many areas of Computer Science (networking, distributed systems, databases, etc.). With the increase in data size and distribution of data, problems...
github.com/yogi1510/Programming-for-Data-Science-with-Python-Udacity-Nanodegree
This repository contains projects did for Udacity Programming For Data Science With Python Nanodegree. (⭐ 2)
github.com/prefuse/Prefuse
Prefuse is a set of software tools for creating rich interactive data visualizations in the Java programming language. Prefuse supports a rich set of features for data modeling, visualization, and interaction. It provides optimized data structures for tables,…
arxiv.org/abs/1807.09754v1
We report development of a data infrastructure for drug repurposing that takes advantage of two currently available chemical ontologies. The data infrastructure includes a database of compound- target associations augmented with molecular ontological...
github.com/roman79/DublinDataEngineering
The Open Source resources in Data Engineering, Machine Learning, Data Science areas, inspired by [The Open-Source Data Science Masters] (http://datasciencemasters.org/). (⭐ 8)
arxiv.org/abs/1709.02261v3
It is common for authors to communicate their results in graphical figures, but those data are frequently unavailable for reanalysis. Reconstructing data points from a figure manually requires the author to measure the coordinates either on printed p...
github.com/darshilparmar/tokyo-olympic-azure-data-engineering-project
tokyo-olympic-azure-data-engineering-project (⭐ 221)
arxiv.org/abs/1905.03040v1
Community detection in graphs, data clustering, and local pattern mining are three mature fields of data mining and machine learning. In recent years, attributed subgraph mining is emerging as a new powerful data mining task in the intersection of th...