Callhome
Emerging29papers using it
2021first seen
Dataset Card for the Callhome dataset for speaker diarization The CALLHOME Corpus is a collection of unscripted telephone conversations between native speakers in Chinese, English, German, Japanese and Spanish. This is a processed version of the original Callhome dataset from the TalkBank corpora taken from here. It co
Papers using Callhome (29)
- MixRep: Hidden Representation Mixup for Low-Resource Speech RecognitionLibriConvo: Simulating Conversations from Read Literature for ASR and DiarizationWhisper Speaker Identification: Leveraging Pre-Trained Multilingual
Transformers for Robust Speaker EmbeddingsSpeaker-Aware Simulation Improves Conversational Speech RecognitionO-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker DiarizationLESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild DataUniversal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity DetectionSEAL: Speaker Error Correction using Acoustic-conditioned Large Language
ModelsTight integration of neural- and clustering-based diarization through
deep unfolding of infinite Gaussian mixture modelImproving Transformer-based End-to-End Speaker Diarization by Assigning
Auxiliary Losses to Attention HeadsMulti-scale Speaker Diarization with Dynamic Scale WeightingTowards Neural Diarization for Unlimited Numbers of Speakers Using
Global and Local AttractorsLow-Latency Speech Separation Guided Diarization for Telephone
ConversationsTarget Speaker Voice Activity Detection with Transformers and Its
Integration with End-to-End Neural DiarizationNeural Diarization with Non-autoregressive Intermediate AttractorsUnified Modeling of Multi-Talker Overlapped Speech Recognition and
Diarization with a Sidecar SeparatorAttention-based Encoder-Decoder End-to-End Neural Diarization with
Embedding EnhancerDecoder-only Architecture for Speech Recognition with CTC Prompts and
Text Data AugmentationDiaPer: End-to-End Neural Diarization with Perceiver-Based AttractorsIntegrating Text Inputs For Training and Adapting RNN Transducer ASR
ModelsGeneration of Speaker Representations Using Heterogeneous Training Batch
AssemblyImproving the Training Recipe for a Robust Conformer-based Hybrid ModelUtterance-by-utterance overlap-aware neural diarization with Graph-PITUSED: Universal Speaker Extraction and DiarizationSemi-Autoregressive Streaming ASR With Label ContextAutomatic Speech Recognition System-Independent Word Error Rate
EstimationLeveraging Speaker Embeddings in End-to-End Neural Diarization for
Two-Speaker ScenariosTowards Unsupervised Speaker Diarization System for Multilingual
Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse
AutoencodersLS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction