arxiv.org/abs/1809.02869v3
Machine teaching addresses the problem of finding the best training data that can guide a learning algorithm to a target model with minimal effort. In conventional settings, a teacher provides data that are consistent with the true data distribution....
www.bing.com/ck/a?!&&p=f056390f5d8847e5ccab1295f9cab1e19be3fe84acbcb7cea6d52c36e82867a6JmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=2eb6d249-1fef-62a8-308f-c5581ed86367&u=a1aHR0cHM6Ly9zdGFja292ZXJmbG93LmNvbS9xdWVzdGlvbnMvMTU5NDM3NjkvaG93LWRvLWktZ2V0LXRoZS1yb3ctY291bnQtb2YtYS1wYW5kYXMtZGF0YWZyYW1l&ntb=1
Apr 11, 2013 · 173 How do I get the row count of a Pandas DataFrame? This table summarises the different situations in which you'd want to count something in a DataFrame (or Series, for …
arxiv.org/abs/1903.12525v1
Bloom Filter is a probabilistic membership data structure and it is excessively used data structure for membership query. Bloom Filter becomes the predominant data structure in approximate membership filtering. Bloom Filter extremely enhances the que...
arxiv.org/abs/1501.01941v3
Bloom filters are probabilistic data structures commonly used for approximate membership problems in many areas of Computer Science (networking, distributed systems, databases, etc.). With the increase in data size and distribution of data, problems...
github.com/yogi1510/Programming-for-Data-Science-with-Python-Udacity-Nanodegree
This repository contains projects did for Udacity Programming For Data Science With Python Nanodegree. (⭐ 2)
github.com/prefuse/Prefuse
Prefuse is a set of software tools for creating rich interactive data visualizations in the Java programming language. Prefuse supports a rich set of features for data modeling, visualization, and interaction. It provides optimized data structures for tables,…
arxiv.org/abs/1807.09754v1
We report development of a data infrastructure for drug repurposing that takes advantage of two currently available chemical ontologies. The data infrastructure includes a database of compound- target associations augmented with molecular ontological...
github.com/roman79/DublinDataEngineering
The Open Source resources in Data Engineering, Machine Learning, Data Science areas, inspired by [The Open-Source Data Science Masters] (http://datasciencemasters.org/). (⭐ 8)
github.com/sfikas/medical-imaging-datasets
A list of Medical imaging datasets. (⭐ 2514)
arxiv.org/abs/1709.02261v3
It is common for authors to communicate their results in graphical figures, but those data are frequently unavailable for reanalysis. Reconstructing data points from a figure manually requires the author to measure the coordinates either on printed p...
github.com/darshilparmar/tokyo-olympic-azure-data-engineering-project
tokyo-olympic-azure-data-engineering-project (⭐ 221)
arxiv.org/abs/1905.03040v1
Community detection in graphs, data clustering, and local pattern mining are three mature fields of data mining and machine learning. In recent years, attributed subgraph mining is emerging as a new powerful data mining task in the intersection of th...
arxiv.org/abs/2001.11324v1
The Symposium on Data Mining and Applications (SDMA 2014) is aimed to gather researchers and application developers from a wide range of data mining related areas such as statistics, computational intelligence, pattern recognition, databases, Big Dat...
github.com/jiqizhixin/Artificial-Intelligence-Terminology-Database
A comprehensive mapping database of English to Chinese technical vocabulary in the artificial intelligence domain (⭐ 2011)
www.bing.com/ck/a?!&&p=0452e621b1543522cd70ed21c7dca03e2732fb74840a98b012112e3279ab8bc7JmltdHM9MTc3MjQ5NjAwMA&ptn=3&ver=2&hsh=4&fclid=16b55049-f68a-6212-20f7-4758f73763a5&u=a1aHR0cHM6Ly93d3cuY2Vuc3VzLmdvdi9kYXRhLmh0bWw&ntb=1
Aug 28, 2025 · Access demographic, economic and population data from the U.S. Census Bureau. Explore census data with visualizations and view tutorials.
arxiv.org/abs/2301.02200v2
Autonomous Driving (AD), the area of robotics with the greatest potential impact on society, has gained a lot of momentum in the last decade. As a result of this, the number of datasets in AD has increased rapidly. Creators and users of datasets can...
arxiv.org/abs/2003.06797v1
In the past years we have witnessed the rise of new data sources for the potential production of official statistics, which, by and large, can be classified as survey, administrative, and digital data. Apart from the differences in their generation a...
github.com/prateekmehta59/Celebrity-Face-Recognition-Dataset
Dataset of around 800k images consisting of 1100 Famous Celebrities and an Unknown class to classify unknown faces (⭐ 158)
github.com/richard512/Little-Big-Data
Data describing topics ranging from Cars and Air Travel to Billionaires and Celebrities (⭐ 75)
en.wikipedia.org/wiki/Data_independence
Data independence is the type of data transparency that matters for a centralized DBMS. It refers to the immunity of user applications to changes made