VoxCeleb
Canonical38papers using it
2021first seen
A speaker-recognition dataset of utterances from thousands of celebrities collected from YouTube.
Papers using VoxCeleb (38)
- Text-Independent Speaker Verification Using Discrete Audio TokensNeural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi SpeechTASLA: Text-Aligned Speech Tokens with Multiple Layer-AggregationNonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-SpeechIsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environmentsRing Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech SeparationVclip: Face-based Speaker Generation by Face-voice Association LearningRethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker VerificationMagnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognitionEffective Modeling of Critical Contextual Information for TDNN-based Speaker VerificationShort-Segment Speaker Verification with Pre-trained Models and Multi-Resolution EncoderAny-to-any Speaker Attribute Perturbation for Asynchronous Voice AnonymizationClustering-based hard negative sampling for supervised contrastive speaker verificationMGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise
Separable Module for Speaker VerificationMulti-Frequency Information Enhanced Channel Attention Module for
Speaker Representation LearningLarge-scale Self-Supervised Speech Representation Learning for Automatic
Speaker VerificationEDITnet: A Lightweight Network for Unsupervised Domain Adaptation in
Speaker VerificationDisentangling Voice and Content with Self-Supervision for Speaker
RecognitionIntroducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering DetectionCAM++: A Fast and Efficient Network for Speaker Verification Using
Context-Aware MaskingA Comparative Study of Modular and Joint Approaches for
Speaker-Attributed ASR on Monaural Long-Form AudioUnsupervised Speech Enhancement with speech recognition embedding and
disentanglement lossesConvolution-Based Channel-Frequency Attention for Text-Independent
Speaker VerificationImproving Text-Independent Speaker Verification with Auxiliary Speakers
Using GraphMulti-query multi-head attention pooling and Inter-topK penalty for
speaker verificationMultiSV: Dataset for Far-Field Multi-Channel Speaker VerificationMFA: TDNN with Multi-scale Frequency-channel Attention for
Text-independent Speaker Verification with Short UtterancesToroidal Probabilistic Spherical Discriminant AnalysisLaugh Betrays You? Learning Robust Speaker Representation From Speech
Containing Non-Verbal FragmentsModel Compression for DNN-based Speaker Verification Using Weight
QuantizationDistance-based Weight Transfer from Near-field to Far-field Speaker
VerificationSelf-FiLM: Conditioning GANs with self-supervised representations for
bandwidth extension based speaker recognitionOrdered and Binary Speaker EmbeddingLeveraging ASR Pretrained Conformers for Speaker Verification through
Transfer Learning and Knowledge DistillationSE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker
RecognitionAV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech
Separation By Leveraging Narrow- and Cross-Band ModelingM-Vec: Matryoshka Speaker Embeddings with Flexible DimensionsNeural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions