English
Emerging33papers using it
2021first seen
Papers using English (33)
- LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic ContextBreaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin ScriptsAccent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram MaskingWord stress in self-supervised speech models: A cross-linguistic comparisonExploring Cross-Lingual Voice Conversion Methods for Anonymizing Low-Resource Text-to-SpeechRethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMsAn Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec DecodingUtterance-Level Methods for Identifying Reliable ASR-Output for Child Speechfindsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and EmbeddingUnsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora OnlyE2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech SynthesisUnsupervised lexicon learning from speech is limited by representations rather than clusteringParallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-SpeechSpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to
Accelerate Your Speech-LLM DevelopmentEvaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning
and Language Identification for Improved Low-resource PerformanceLeveraging supplementary text data to kick-start automatic speech
recognition system development with limited transcriptionsEditSpeech: A Text Based Speech Editing System Using Partial Inference
and Bidirectional FusionAnalysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAEAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised
Speech ModelsTTS-Guided Training for Accent Conversion Without Parallel DataSyllable Discovery and Cross-Lingual Generalization in a Visually
Grounded, Self-Supervised Speech ModelParaformer-v2: An improved non-autoregressive transformer for
noise-robust speech recognitionAdvocating Character Error Rate for Multilingual ASR EvaluationSTTATTS: Unified Speech-To-Text And Text-To-Speech ModelIntent Classification Using Pre-trained Language Agnostic Embeddings For
Low Resource LanguagesVisually Grounded Keyword Detection and Localisation for Low-Resource
LanguagesLearning Multilingual Expressive Speech Representation for Prosody
Prediction without Parallel DataZero Resource Cross-Lingual Part Of Speech TaggingContextualized Automatic Speech Recognition with Dynamic VocabularyVECL-TTS: Voice identity and Emotional style controllable Cross-Lingual
Text-to-SpeechEnhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in
Any-to-One Voice ConversionTextless NLP -- Zero Resource Challenge with Low Resource ComputeLightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech