Improved Robustness To Disfluencies In Rnn-transducer Based Speech Recognition
2020 Β· Valentin Mendelev, Tina Raissi, Guglielmo Camporese, et al.
Abstract
Automatic Speech Recognition (ASR) based on Recurrent Neural Network Transducers (RNN-T) is gaining interest in the speech community. We investigate data selection and preparation choices aiming for improved robustness of RNN-T ASR to speech disfluencies with a focus on partial words. For evaluation we use clean data, data with disfluencies and a separate dataset with speech affected by stuttering. We show that after including a small amount of data with disfluencies in the training set the recognition accuracy on the tests with disfluencies and stuttering improves. Increasing the amount of training data with disfluencies gives additional gains without degradation on the clean data. We also show that replacing partial words with a dedicated token helps to get even better accuracy on utterances with disfluencies and stutter. The evaluation of our best model shows 22.5% and 16.4% relative WER reduction on those two evaluation sets.
Authors
(none)
Tags
Stats
Related papers
- Improving RNN Transducer Modeling For End-to-end Speech Recognition (2019)0.00
- Segaug: Ctc-aligned Segmented Augmentation For Robust Rnn-transducer Based Speech Recognition (2025)3.58
- Exploring Architectures, Data And Units For Streaming End-to-end Speech Recognition With Rnn-transducer (2018)16.21
- Alignment Restricted Streaming Recurrent Neural Network Transducer (2020)11.19
- Integrating Text Inputs For Training And Adapting RNN Transducer ASR Models (2022)9.59
- Efficient Minimum Word Error Rate Training Of Rnn-transducer For End-to-end Speech Recognition (2020)11.19
- Improved Neural Language Model Fusion For Streaming Recurrent Neural Network Transducer (2020)8.82
- A Comparison Of Streaming Models And Data Augmentation Methods For Robust Speech Recognition (2021)2.26