TIMIT
Canonical47papers using it
2021first seen
A phonetically-transcribed read-speech corpus widely used for acoustic-phonetic and ASR research.
Papers using TIMIT (47)
- AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition BaselinesGradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMsMultilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic ProgrammingFlowW2N: Whispered-to-Normal Speech Conversion via Flow-MatchingSingle Channel Blind Dereverberation of Speech SignalsBFA: Real-time Multilingual Text-to-speech Forced AlignmentEvaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCsState-Space Models in Efficient Whispered and Multi-dialect Speech RecognitionA Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal TransportBack To Supervision: Boosting Word Boundary Detection Through Frame ClassificationImproving Whispered Speech Recognition Performance using
Pseudo-whispered based Data AugmentationNormalizing Flow based Hidden Markov Models for Classification of Speech
Phones with ExplainabilityComplex Recurrent Variational Autoencoder with Application to Speech
EnhancementFoster Strengths and Circumvent Weaknesses: a Speech Enhancement
Framework with Two-branch Collaborative LearningRepresentative Subset Selection for Efficient Fine-Tuning in
Self-Supervised Speech RecognitionRepresentation Learning With Hidden Unit Clustering For Low Resource
Speech ApplicationsUnsupervised Domain Adaptation in Speech Recognition using Phonetic
FeaturesTime-frequency Network for Robust Speaker RecognitionImproving Deep Attractor Network by BGRU and GMM for Speech SeparationREBORN: Reinforcement-Learned Boundary Segmentation with Iterative
Training for Unsupervised ASRTradition or Innovation: A Comparison of Modern ASR Methods for Forced
AlignmentArtificial bandwidth extension using deep neural network and $H^\infty$
sampled-data control theoryUnsupervised Speech Segmentation and Variable Rate Representation
Learning using Segmental Contrastive Predictive CodingNoisy Speech Based Temporal Decomposition to Improve Fundamental
Frequency EstimationEstimation of speaker age and height from speech signal using bi-encoder
transformer mixture modelNeuraGen-A Low-Resource Neural Network based approach for Gender
ClassificationRobust Disentangled Variational Speech Representation Learning for
Zero-shot Voice ConversionLearning Phone Recognition from Unpaired Audio and Phone Sequences Based
on Generative Adversarial NetworkAn Experimental Study on Private Aggregation of Teacher Ensemble
Learning for End-to-End Speech RecognitionPhoneme Segmentation Using Self-Supervised Speech ModelsEURO: ESPnet Unsupervised ASR Open-source ToolkitEnhancing Unsupervised Speech Recognition with Diffusion GANsWeakly-supervised forced alignment of disfluent speech using
phoneme-level modelingTimestamped Embedding-Matching Acoustic-to-Word CTC ASRPDPCRN: Parallel Dual-Path CRN with Bi-directional Inter-Branch
Interactions for Multi-Channel Speech EnhancementHierarchical Modeling of Spatial Cues via Spherical Harmonics for
Multi-Channel Speech EnhancementEfficient Multi-Channel Speech Enhancement with Spherical Harmonics
Injection for Directional EncodingUnsupervised Speech Recognition with N-Skipgram and Positional Unigram
MatchingPhasePerturbation: Speech Data Augmentation via Phase Perturbation for
Automatic Speech RecognitionOn Speech Pre-emphasis as a Simple and Inexpensive Method to Boost
Speech EnhancementLeveraging Self-Supervised Models for Automatic Whispered Speech
RecognitionMaskCycleGAN-based Whisper to Normal Speech ConversionBack to Supervision: Boosting Word Boundary Detection through Frame
ClassificationBEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken
Term DetectionDomain Adaptation and Autoencoder Based Unsupervised Speech EnhancementDisentangled Speech Representation Learning Based on Factorized
Hierarchical Variational Autoencoder with Self-Supervised ObjectiveQuartered Spectral Envelope and 1D-CNN-based Classification of Normally
Phonated and Whispered Speech