π Datasets β Awesome Federated Learning
Loading datasetsβ¦
Loading datasetsβ¦
1,065 datasets & benchmarks β 17 canonical foundations plus emerging datasets mined from recent papers. Each links to the papers that use it.
60,000 32Γ32 color images in 10 classes β a small, standard image-classification benchmark.
70,000 28Γ28 grayscale images of handwritten digits (0β9) β the classic image-classification benchmark.
Like CIFAR-10 but with 100 fine classes (grouped into 20 superclasses), 600 images each.
A drop-in MNIST replacement with 70,000 grayscale images across 10 clothing categories.
Dataset Card for FEMNIST The FEMNIST dataset is a part of the LEAF benchmark. It represents image classification of handwritten digits, lower and uppercase letters, giving 62 unique labels. Dataset Details Dataset Description Each sample is comprised of a (28x28) grayscale image, writer_id, hsf_id, and character. Curated by: LEAF License: BSD 2-Clause License Dataset Sources The FEMNIST is a preprocessed (in a way that resembles preprocessing for⦠See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/femnist.
EMNIST is a dataset that contains handwritten character samples and is used to evaluate the performance of machine learning models in recognizing and classifying these characters.
'Tiny ImageNet' is a dataset that contains 200 classes of images, each with 500 training images, used to evaluate performance in image classification tasks.
F-MNIST (Fashion-MNIST) is a dataset that contains grayscale images of clothing items and is used to evaluate the performance of machine learning models, particularly in the context of backdoor attacks in federated learning.
Dataset Card for Street View House Numbers Dataset Summary SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and formatting. It can be seen as similar in flavor to MNIST (e.g., the images are of small cropped digits), but incorporates an order of magnitude more labeled data (over 600,000 digit images) and comes from a significantly harder, unsolved, real world problem⦠See the full description on the dataset page: https://huggingface.co/datasets/ufldl-stanford/svhn.
ImageNet is a large-scale dataset containing millions of labeled images across thousands of categories, commonly used to evaluate the performance of image classification algorithms.
200k celebrity face images annotated with 40 binary attributes and facial landmarks.
The 'Shakespeare' dataset is a benchmarking dataset used to evaluate the performance of federated learning algorithms in terms of accuracy, convergence time, communication overhead, energy consumption, and robustness to non-IID data.
The 'ToN-IoT' dataset is a benchmark that contains network traffic data specifically designed for evaluating intrusion detection systems in Internet of Things (IoT) environments.
The 'UNSW-NB-15' dataset is a benchmark that contains network traffic data used to evaluate intrusion detection systems, particularly in the context of identifying various types of cyber attacks.
CINIC-10 is a dataset that contains images for evaluating machine learning models, specifically designed to benchmark performance in image classification tasks.
Dataset Card for PACS PACS is an image dataset for domain generalization. It consists of four domains, namely Photo (1,670 images), Art Painting (2,048 images), Cartoon (2,344 images), and Sketch (3,929 images). Each domain contains seven categories (labels): Dog, Elephant, Giraffe, Guitar, Horse, and Person. The total number of sample is 9991. Dataset Details PACS DG dataset is created by intersecting the classes found in Caltech256 (Photo), Sketchy (Photo, Sketch)β¦ See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/pacs.
Data downloaded from WILDS (Download, paper, project). This dataset contains some copyrighted material whose use has not been specifically authorized by the copyright owners. In an effort to advance scientific research, we make this material available for academic research. We believe this constitutes a fair use of any such copyrighted material as provided for in section 107 of the US Copyright Law. In accordance with Title 17 U.S.C. Section 107, the material on this site is distributed⦠See the full description on the dataset page: https://huggingface.co/datasets/wltjr1007/DomainNet.
Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained analysis of system⦠See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.
MIMIC-IV is a publicly available critical care database that contains de-identified health data from ICU patients, used to evaluate predictive models for early sepsis detection.
The 'Reddit' dataset is used to evaluate the effectiveness of personalized federated learning approaches by providing a collection of user-generated content for assessing model performance in a collaborative learning environment.
The 'chest X-ray' dataset contains 112,120 chest X-ray images used to evaluate the diagnosis of various diseases in a medical imaging context.
LEAF is a benchmark dataset used to evaluate federated learning algorithms, containing various tasks designed to simulate the challenges of training machine learning models across distributed networks of mobile devices.
MovieLens is a dataset used to evaluate recommendation systems, containing user preferences for movies.
AG News is a benchmark dataset that contains news articles categorized into four classes, used to evaluate text classification models in the context of federated learning.
Raw network data was collected over a period of 5 days, Monday through Friday, and stored in PCAP files. Monday was used to create most of the Benign data, while the Attack-Network implemented various types of attacks over the next 4 days, such as Brute Force connections (FTP and SSH), several types of DoS attacks, as well as a Botnet attack, Infiltration attacks and subsequent Port-Scanning activity. The PCAP data was processed using a tool developed by one of the authors of [1], called⦠See the full description on the dataset page: https://huggingface.co/datasets/bvk/CICIDS-2017.
Stack Overflow is a dataset used to evaluate the performance of machine learning models, particularly in the context of personalized federated learning, by providing a platform for analyzing user-generated content and interactions.
CARLA is a benchmark and simulator used to evaluate autonomous vehicles' interactions with human drivers in diverse geographic areas, focusing on trajectory forecasting and human-robot interactions.
CIFAR is a benchmark dataset that contains a collection of images used to evaluate the performance of machine learning models, particularly in the context of image classification tasks.
The 'Edge-IIoTset' is a dataset used to evaluate intrusion detection performance in heterogeneous Internet of Things (IoT) networks.
Dataset Card for "imdb" Dataset Summary Large Movie Review Dataset. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure⦠See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/imdb.
The N-BaIoT dataset is a benchmark for evaluating intrusion detection systems in heterogeneous Internet of Things (IoT) networks, containing diverse real-world scenarios related to IoT security.
The UCI-HAR dataset is a benchmark that contains human activity recognition data collected from smartphones, used to evaluate machine learning models in the context of federated learning.
The Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset contains clinical, imaging, genetic, and biospecimen data aimed at evaluating the progression of Alzheimer's disease and related disorders.
Data Source https://www.kaggle.com/datasets/paultimothymooney/chest-xray-pneumonia Dataset Card Authors Mahadi Hassan Dataset Card Contact mahadise01@gmail.com Linkdin: https://www.linkedin.com/in/mahadise01 Github: https://github.com/Mahadih534
CheXpert is a public dataset containing chest X-ray images used to evaluate the performance of deep learning models in detecting various medical findings.
CIFAR-10-LT is a long-tailed version of the CIFAR-10 dataset used to evaluate federated learning methods on non-IID and imbalanced data distributions.
Dataset Card for digits dataset Optical recognition of handwritten digits dataset Note - How to load this dataset directly with the datasets library from datasets import load_dataset dataset = load_dataset("sklearn-docs/digits",header=None) Dataset Summary This is a copy of the test set of the UCI ML hand-written digits datasets https://archive.ics.uci.edu/ml/datasets/Optical+Recognition+of+Handwritten+Digits The data set contains images of hand-written⦠See the full description on the dataset page: https://huggingface.co/datasets/sklearn-docs/digits.
The 'eICU' dataset contains clinical data from intensive care unit patients and is used to evaluate predictive models for early sepsis detection.
FLAIR is a large-scale annotated image dataset containing 429,078 images from 51,414 Flickr users, designed to evaluate multi-label classification in the context of federated learning, capturing challenges like heterogeneous user data and long-tailed label distribution.
Human Activity Recognition (HAR) is a dataset/benchmark used to evaluate models that identify and classify human activities based on data collected from various sensors.
ISIC-2019 is a benchmark dataset that contains a collection of dermoscopic images used to evaluate the performance of algorithms in the context of skin lesion classification.
Dataset Card for Kitti The Kitti dataset. The Kitti object detection and object orientation estimation benchmark consists of 7481 training images and 7518 test images, comprising a total of 80.256 labeled objects
Language modeling resources to be used in conjunction with the LibriSpeech ASR corpus.
MIMIC-CXR is a dataset that contains chest X-ray images and associated clinical reports, used to evaluate longitudinal medical report generation and the effectiveness of federated learning methods in capturing temporal dynamics in patient data.
MIMIC-III is a clinical benchmark dataset that contains de-identified health data for patients and is used to evaluate the reliability and performance of federated learning in a medical context.
The PAMAP-2 dataset is a benchmark that contains data from various physical activities collected using multiple sensors, and it is used to evaluate algorithms for activity recognition and classification in the context of wearable computing.
The 'Sent140' dataset contains a collection of tweets labeled for sentiment analysis and is used to evaluate the performance of models in understanding and predicting sentiment in text data.
The 'Sentiment-140' dataset contains 1.6 million tweets labeled for sentiment analysis, and it is used to evaluate the performance of sentiment classification models.
CIFAR-100-LT is a long-tailed version of the CIFAR-100 dataset used to evaluate federated learning methods on non-IID and imbalanced data distributions.
The 'COVID-19' dataset is a clinical dataset used to evaluate the performance of decentralized machine learning methods, particularly in the context of non-IID data challenges.
Dataset Card for GTSRB Dataset Summary The German Traffic Sign Benchmark is a multi-class, single-image classification challenge held at the International Joint Conference on Neural Networks (IJCNN) 2011. We cordially invite researchers from relevant fields to participate: The competition is designed to allow for participation without special domain knowledge. Our benchmark has the following properties: Single-image, multi-class classification problem More than 40β¦ See the full description on the dataset page: https://huggingface.co/datasets/bazyl/GTSRB.
Dataset Card for PathMedMNIST This is currently the PathMedMNIST part of the MedMNIST dataset in 64x64 resolution. Dataset Source Website: https://medmnist.com/ The entire dataset on HF albertvillanova/medmnist-v2
MuJoCo is a benchmark that contains a set of reinforcement learning tasks used to evaluate algorithms in the context of reinforcement learning from human feedback.
The 'Office-10' dataset is a benchmark that contains images from ten different categories across three distinct domains (Amazon, Webcam, and DSLR) and is used to evaluate domain adaptation and transfer learning methods in the context of heterogeneous feature distributions.
PathMNIST is a dataset that contains images of handwritten digits derived from the MNIST dataset, used to evaluate the performance of federated learning under varying levels of client heterogeneity.
The 'ROAD' dataset is used to evaluate network anomaly detection in the context of federated learning.