GhostSecurityTeamIO/OnlyFans-DataScraper
OnlyFans DataScraper (Python 3.9.X) (⭐ 35)
OnlyFans DataScraper (Python 3.9.X) (⭐ 35)
Generative modeling has been used frequently in synthetic data generation. Fairness and privacy are two big concerns for synthetic data. Although Recent GAN [\cite{goodfellow2014generative}] based methods show good results in preserving privacy, the...
Today, the operating TAIGA (Tunka Advanced Instrument for cosmic rays and Gamma Astronomy) experiment continuously produces and accumulates a large volume of raw astroparticle data. To be available for the scientific community these data should be we...
This article describes the use of metadata and standards in the Social Impact Data Commons to expose official statisticians to an innovative project built on actionable and evaluable metadata, which produces a FAIR data system. We begin by introducin...
OpenMetadata is a unified metadata platform for data discovery, data observability, and data governance powered by a central metadata repository, in-depth column level lineage, and seamless team collaboration. (⭐ 8859)
What is metadata? Metadata is information—such as author, creation date or file size—that describes a data point or data set. Metadata can improve a data system’s functions and make it easier to search …
about subject descriptions of data and token codes for the data. We also have statements in a meta language describing the data relationships and transformations
I was going through the aws-certified-database-specialty-dbs course in Udemy by Riaze and Stephane Maarek for AWS Database specialty Certification. There's a very confusing section (DMS Statistics a...
Visual Genome is a dataset connecting structured image information with English language. We present ``Hindi Visual Genome'', a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task and multim...
Most datasets of interest to the analytics industry are impacted by various forms of human bias. The outcomes of Data Analytics [DA] or Machine Learning [ML] on such data are therefore prone to replicating the bias. As a result, a large number of bia...
The new generation of cloud data warehouses (CDWs) brings large amounts of data and compute power closer to users in enterprises. The ability to directly access the warehouse data, interactively analyze and explore it at scale can empower users to im...
We present PyTorch Frame, a PyTorch-based framework for deep learning over multi-modal tabular data. PyTorch Frame makes tabular deep learning easy by providing a PyTorch-based data structure to handle complex tabular data, introducing a model abstra...
Wireless data aggregation (WDA), referring to aggregating data distributed at devices (e.g., sensors and smartphone), is a common operation in 5G-and-beyond machine-type communications to support Internet-of-Things (IoT), which lays the foundation fo...
In the past few years, we have envisioned an increasing number of businesses start driving by big data analytics, such as Amazon recommendations and Google Advertisements. At the back-end side, the businesses are powered by big data processing platfo...
List of image datasets with any kind of litter, garbage, waste and trash (⭐ 332)
Use of medical data, also known as electronic health records, in research helps develop and advance medical science. However, protecting patient confidentiality and identity while using medical data for analysis is crucial. Medical data can be in the...
These recommendations are the result of reflections by scientists and experts who are, or have been, involved in the preservation of high-energy physics data. The work has been done under the umbrella of the Data Lifecycle panel of the International...
A collection of Rfast2 functions for data analysis. Note 1: The vast majority of the functions accept matrices only, not data.frames. Note 2: Do not have matrices or vectors with have missing data (i.e NAs). We do no check about them and C++ internally transfo…
US Senate financial reports and structured data (⭐ 12)
Randomized evaluations of educational technology produce log data as a bi-product: highly granular data student and teacher usage. These datasets could shed light on causal mechanisms, effect heterogeneity, or optimal use. However, there are methodol...