Iterative Shallow Fusion Of Backward Language Model For End-to-end Speech Recognition
2023 Β· Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, et al.
Abstract
We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescoring ASR hypotheses as post-processing. In the proposed SF, we iteratively apply the BLM to partial ASR hypotheses in the backward direction (i.e., from the possible next token to the start symbol) during decoding, substituting the newly calculated BLM scores for the scores calculated at the last iteration. To enhance the effectiveness of this iterative SF (ISF), we train a partial sentence-aware BLM (PBLM) using reversed text data including partial sentences, considering the framework of ISF. In experiments using an attention-based encoder-decoder ASR system, we confirmed that ISF using the PBLM shows comparable performance with SF using the FLM. By performing ISF, early pruning of prospective
Authors
(none)
Tags
Stats
Related papers
- Delayed Fusion: Integrating Large Language Models Into First-pass Decoding In End-to-end Speech Recognition (2025)5.84
- Transfer Learning Of Language-independent End-to-end ASR With Language Model Fusion (2018)0.00
- An Analysis Of Incorporating An External Language Model Into A Sequence-to-sequence Model (2017)16.25
- Let's Fuse Step By Step: A Generative Fusion Decoding Algorithm With Llms For Robust And Instruction-aware ASR And OCR (2024)0.00
- Adapting Speech Foundation Models For Unified Multimodal Speech Recognition With Large Language Models (2025)0.00
- Multilingual And Fully Non-autoregressive ASR With Large Language Model Fusion: A Comprehensive Study (2024)0.00
- Internal Language Model Estimation Based Language Model Fusion For Cross-domain Code-switching Speech Recognition (2022)0.00
- Improved Neural Language Model Fusion For Streaming Recurrent Neural Network Transducer (2020)8.82