English datasets
Emerging33papers using it
2021first seen
The 'English datasets' benchmark contains various datasets used to evaluate the performance and security of speech synthesis systems, particularly in the context of adversarial attacks and defenses.
Papers using English datasets (33)
- LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic ContextBreaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin ScriptsAccent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram MaskingWord stress in self-supervised speech models: A cross-linguistic comparisonExploring Cross-Lingual Voice Conversion Methods for Anonymizing Low-Resource Text-to-SpeechRethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMsAn Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec DecodingUtterance-Level Methods for Identifying Reliable ASR-Output for Child Speechfindsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and EmbeddingUnsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora OnlyE2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech SynthesisUnsupervised lexicon learning from speech is limited by representations rather than clusteringParallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-SpeechSpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM DevelopmentEvaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource PerformanceLeveraging supplementary text data to kick-start automatic speech recognition system development with limited transcriptionsEditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional FusionAnalysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAEAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised Speech ModelsTTS-Guided Training for Accent Conversion Without Parallel DataSyllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech ModelParaformer-v2: An improved non-autoregressive transformer for noise-robust speech recognitionAdvocating Character Error Rate for Multilingual ASR EvaluationSTTATTS: Unified Speech-To-Text And Text-To-Speech ModelIntent Classification Using Pre-trained Language Agnostic Embeddings For
Low Resource LanguagesVisually Grounded Keyword Detection and Localisation for Low-Resource
LanguagesLearning Multilingual Expressive Speech Representation for Prosody
Prediction without Parallel DataZero Resource Cross-Lingual Part Of Speech TaggingContextualized Automatic Speech Recognition with Dynamic VocabularyVECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-SpeechEnhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice ConversionTextless NLP -- Zero Resource Challenge with Low Resource ComputeLightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech