📊 Datasets — Awesome Computer Vision
Loading datasets…
Loading datasets…
1,264 datasets & benchmarks — 34 canonical foundations plus emerging datasets mined from recent papers. Each links to the papers that use it.
Common Objects in Context — 330k images with object-detection, segmentation, keypoint, and captioning annotations.
~1.28M labeled images across 1,000 categories (ILSVRC) — the standard large-scale image-classification benchmark.
Urban street scenes with fine pixel-level semantic segmentation across 30 classes, for driving perception.
An object-detection and segmentation benchmark of 20 object categories from the PASCAL Visual Object Classes challenge.
A scene-parsing dataset with dense pixel annotations over 150 semantic categories.
Autonomous-driving benchmarks (stereo, optical flow, detection, tracking) recorded from a car in urban traffic.
Dataset Card for HICO-DET Dataset Dataset Summary HICO-DET is a dataset for detecting human-object interactions (HOI) in images. It contains 47,776 images (38,118 in train set and 9,658 in test set), 600 HOI categories constructed by 80 object categories and 117 verb classes. HICO-DET provides more than 150k annotated human-object pairs. V-COCO provides 10,346 images (2,533 for training, 2,867 for validating and 4,946 for testing) and 16,199 person instances. Each person… See the full description on the dataset page: https://huggingface.co/datasets/zhimeng/hico_det.
Pascal VOC 2012 is a dataset used for evaluating semantic segmentation algorithms, containing annotated images for object detection and classification tasks.
DAVIS-17 is a benchmark dataset used to evaluate Zero-Shot Video Object Segmentation (ZS-VOS) performance, containing a collection of video sequences with annotated object instances.
The V-COCO dataset is a benchmark that contains images annotated with human-object interactions, used to evaluate the performance of models in detecting and classifying these interactions.
YouTube-VOS is a benchmark dataset used to evaluate semi-supervised video object segmentation by providing a collection of video sequences with annotated object masks.
Like CIFAR-10 but with 100 fine classes (grouped into 20 superclasses), 600 images each.
60,000 32×32 color images in 10 classes — a small, standard image-classification benchmark.
A large autonomous-driving dataset with 360° camera, lidar, and radar across 1,000 driving scenes.
A densely-annotated video object-segmentation benchmark (DAVIS 2016/2017).
The MOT-17 dataset is a benchmark for evaluating multi-object tracking algorithms, containing video sequences with annotated object trajectories and identities, specifically designed to assess performance in scenarios with challenges like occlusion.
This is ScanNet dataset containing indoor scenes which is used for 3d object detection, 3d segmentation, etc. Acknowledgement @inproceedings{dai2017scannet, title={ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes}, author={Dai, Angela and Chang, Angel X. and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie{\ss}ner, Matthias}, booktitle = {Proc. Computer Vision and Pattern Recognition (CVPR), IEEE}, year = {2017} }
Large-vocabulary instance segmentation over 1,000+ categories with a long-tailed distribution.
The 'DAVIS'16' dataset is a benchmark used for evaluating video object segmentation algorithms, containing high-resolution video sequences with pixel-level annotations for object instances.
The 'CUB-200-2011' dataset is a benchmark that contains images of 200 bird species, annotated with attributes, and is used to evaluate visual recognition and retrieval systems.
Dataset Card for DanceTrack DanceTrack is a multi-human tracking dataset with two emphasized properties, (1) uniform appearance: humans are in highly similar and almost undistinguished appearance, (2) diverse motion: humans are in complicated motion pattern and their relative positions exchange frequently. We expect the combination of uniform appearance and complicated motion pattern makes DanceTrack a platform to encourage more comprehensive and intelligent multi-object tracking… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/DanceTrack.
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
WIDER FACE dataset is a face detection benchmark dataset, of which images are selected from the publicly available WIDER dataset. We choose 32,203 images and label 393,703 faces with a high degree of variability in scale, pose and occlusion as depicted in the sample images. WIDER FACE dataset is organized based on 61 event classes. For each event class, we randomly select 40%/10%/50% data as training, validation and testing sets. We adopt the same evaluation metric employed in the PASCAL VOC dataset. Similar to MALF and Caltech datasets, we do not release bounding box ground truth for the test images. Users are required to submit final prediction files, which we shall proceed to evaluate.
MOT-20 is a benchmark dataset used to evaluate multiple object tracking methods, containing sequences of video data with annotated object identities to assess tracking performance.
The 'Pascal-5$^i$' dataset is a benchmark used to evaluate few-shot object detection methods by providing a set of images with annotated objects across multiple categories, allowing for the assessment of model performance in recognizing new object classes with limited training examples.
THUMOS-14 is a benchmark dataset that contains untrimmed videos annotated with action boundaries and categories, used to evaluate Temporal Action Localization (TAL) models.
13,320 videos across 101 human-action categories — a standard action-recognition benchmark.
GTEA is composed of 50 recorded videos of 25 participants making two different mixed salads. The videos are captured by a camera with a top-down view onto the work-surface. The participants are provided with recipe steps which are randomly sampled from a statistical recipe model.
General information The overall ACDC dataset was created from real clinical exams acquired at the University Hospital of Dijon. Acquired data were fully anonymized and handled within the regulations set by the local ethical committee of the Hospital of Dijon (France). Our dataset covers several well-defined pathologies with enough cases to (1) properly train machine learning methods and (2) clearly assess the variations of the main physiological parameters obtained from cine-MRI (in… See the full description on the dataset page: https://huggingface.co/datasets/msepulvedagodoy/acdc.
ImageNetVID dataset Usage Please follow the command to use: cat ILSVRC2015_VID.tar.gz.a* > ILSVRC2015_VID.tar.gz cat ILSVRC2017_DET.tar.gz.a* > ILSVRC2017_DET.tar.gz
The 'Pascal Context' dataset is a benchmark that contains annotated images for evaluating multi-task dense prediction tasks in computer vision.
'VBench' is a benchmark used to evaluate the performance of sparse attention mechanisms in video generation tasks.
The MPII dataset is a benchmark for human pose estimation that contains annotated images of people in various activities and is used to evaluate the accuracy of pose estimation methods.
PASCAL VOC 2007 is a benchmark dataset that contains images annotated for object detection tasks, used to evaluate the performance of models in identifying and localizing objects within those images.
GTEA is composed of 50 recorded videos of 25 participants making two different mixed salads. The videos are captured by a camera with a top-down view onto the work-surface. The participants are provided with recipe steps which are randomly sampled from a statistical recipe model.
WILDTRACK is a benchmark dataset used to evaluate multi-camera tracking performance, specifically measuring metrics such as IDF1, MOTA, and MOTP in real-time multi-view 3D tracking systems.
A large-scale human-action video dataset (400/600/700 classes) for action recognition.
BDD-100K is a dataset used to evaluate object detection and semantic segmentation tasks in computer vision, containing a diverse set of images with varying levels of annotation.
The 'COCO-20$^i$' dataset/benchmark contains a subset of the COCO dataset specifically designed for evaluating few-shot object detection methods across 20 categories.
The 'ETH-3D' dataset is a benchmark that contains calibrated images used to evaluate multi-view stereo (MVS) methods for 3D geometry reconstruction.
The 'Human-3.6M' dataset is a benchmark that contains 3D human pose data captured from various activities, used to evaluate algorithms for 3D human pose estimation from monocular videos.
The MARS dataset is a benchmark for evaluating video-based person re-identification methods, containing video sequences with annotated identities to assess the effectiveness of algorithms in recognizing individuals across different camera views.
The 'REAL-275' dataset is a benchmark that contains a diverse set of objects and scenes used to evaluate open-vocabulary 6D object pose estimation methods.
The VOC dataset, or Visual Object Classes dataset, contains images annotated with object classes and is used to evaluate the performance of visual recognition models in detecting and classifying objects.
70,000 28×28 grayscale images of handwritten digits (0–9) — the classic image-classification benchmark.
Code snippet to visualise the position of the box import matplotlib.image as img import matplotlib.pyplot as plt from datasets import load_dataset from matplotlib.patches import Rectangle # Load dataset ds_name = "SaulLu/Stanford-Cars" ds = load_dataset(ds_name, use_auth_token=True) # Extract information for the sample we want to show index = 100 sample = ds["train"][index] box_coord = sample["bbox"][0] img_path = sample["image"].filename # Create plot # define Matplotlib figure and… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/Stanford-Cars.
SemanticKITTI is a dataset that contains 3D point cloud data with annotated semantic labels, used to evaluate the performance of algorithms in semantic segmentation and scene understanding tasks in outdoor environments.
FDDB is a benchmark dataset used to evaluate face detection algorithms, containing a collection of annotated images with varying face orientations and scales.
'ImageNet-100' is a subset of the ImageNet dataset that contains 100 classes of images, used to evaluate the performance of deep neural networks in computer vision tasks.
The iLIDS-VID dataset is a benchmark for evaluating video-based person re-identification, containing videos of individuals captured by different non-overlapping cameras.
The SUN RGB-D dataset is a benchmark that contains RGB-D images used to evaluate semantic segmentation performance in scenarios where depth information may be missing or degraded.
The Synapse dataset is a benchmark used for evaluating medical image segmentation methods, containing annotated medical images for this purpose.
The FER-2013 dataset is a benchmark that contains facial expression images used to evaluate automatic facial expression recognition systems.
Dataset Description, Collection, and Source The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses. In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.
LINEMOD is a benchmark dataset used to evaluate six-degree-of-freedom (6-DOF) object pose estimation algorithms, containing various objects in different poses and occlusions.
The Something-Something dataset (version 2) is a collection of 220,847 labeled video clips of humans performing pre-defined, basic actions with everyday objects. It is designed to train machine learning models in fine-grained understanding of human hand gestures like putting something into something, turning something upside down and covering something with something.
GTEA is composed of 50 recorded videos of 25 participants making two different mixed salads. The videos are captured by a camera with a top-down view onto the work-surface. The participants are provided with recipe steps which are randomly sampled from a statistical recipe model.
The dataset contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes. The (main) aircraft in each image is annotated with a tight bounding box and a hierarchical airplane model label. Aircraft models are organized in a four-levels hierarchy. The four levels, from finer to coarser, are: Model, e.g. Boeing 737-76J. Since certain models are nearly visually indistinguishable, this level is not used in the evaluation. Variant, e.g. Boeing 737-700. A variant collapses all the models that are visually indistinguishable into one class. The dataset comprises 102 different variants. Family, e.g. Boeing 737. The dataset comprises 70 different families. Manufacturer, e.g. Boeing. The dataset comprises 41 different manufacturers. The data is divided into three equally-sized training, validation and test subsets. The first two sets can be used for development, and the latter should be used for final evaluation only.