VoxCeleb2
Canonical20papers using it
2021first seen
VoxCeleb2 Dataset This is the VoxCeleb2 dataset, a large-scale speaker identification dataset. Dataset Description VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Files vox2_dev_mp4_part*: Multipart archive containing MP4 video files vox2_dev_txt: Text file
Papers using VoxCeleb2 (20)
- Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent SpaceCL-UZH submission to the NIST SRE 2024 Speaker Recognition EvaluationFew-Shot Speaker Identification Using Lightweight Prototypical Network
with Feature Grouping and InteractionMulti-View Self-Attention Based Transformer for Speaker RecognitionCan large-scale vocoded spoofed data improve speech spoofing
countermeasure with a self-supervised front end?Self-Supervised Training of Speaker Encoder with Multi-Modal Diverse
Positive PairsSpeaker Recognition Using Isomorphic Graph Attention Network Based
Pooling on Self-Supervised RepresentationSeeing Through the Conversation: Audio-Visual Speech Separation based on
Diffusion ModelTarget Speech Diarization with Multimodal PromptsMechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech SynthesisMulti-task Voice Activated Framework using Self-supervised LearningMore than Words: In-the-Wild Visually-Driven Prosody for Text-to-SpeechLipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction
and Lip ReadingAuto-AVSR: Audio-Visual Speech Recognition with Automatic LabelsA vector quantized masked autoencoder for speech emotion recognitionOne-Step Knowledge Distillation and Fine-Tuning in Using Large
Pre-Trained Self-Supervised Learning Models for Speaker VerificationSpeaker verification using attentive multi-scale convolutional recurrent
networkTarget Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover
StrategyMultilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast
ConformerLibri2Vox Dataset: Target Speaker Extraction with Diverse Speaker
Conditions and Synthetic Data