Residual Adapters For Parameter-efficient ASR Adaptation To Atypical And Accented Speech
2021 Β· Katrin Tomanek, Vicky Zayats, Dirk Padfield, et al.
Abstract
Automatic Speech Recognition (ASR) systems are often optimized to work best for speakers with canonical speech patterns. Unfortunately, these systems perform poorly when tested on atypical speech and heavily accented speech. It has previously been shown that personalization through model fine-tuning substantially improves performance. However, maintaining such large models per speaker is costly and difficult to scale. We show that by adding a relatively small number of extra parameters to the encoder layers via so-called residual adapter, we can achieve similar adaptation gains compared to model fine-tuning, while only updating a tiny fraction (less than 0.5%) of the model parameters. We demonstrate this on two speech adaptation tasks (atypical and accented speech) and for two state-of-the-art ASR architectures.
Authors
(none)
Tags
Stats
Related papers
- Updating Only Encoders Prevents Catastrophic Forgetting Of End-to-end ASR Models (2022)5.24
- ADAPTERMIX: Exploring The Efficacy Of Mixture Of Adapters For Low-resource TTS Adaptation (2023)6.34
- Parameter-efficient Adaptation Of Multilingual Multimodal Models For Low-resource ASR (2024)2.26
- Hyper-parameter Adaptation Of Conformer ASR Systems For Elderly And Dysarthric Speech Recognition (2023)0.00
- Resource-efficient Adaptation Of Speech Foundation Models For Multi-speaker ASR (2024)3.58
- Elp-adapters: Parameter Efficient Adapter Tuning For Various Speech Processing Tasks (2024)7.81
- Parameter-efficient Dysarthric Speech Recognition Using Adapter Fusion And Householder Transformation (2023)0.00
- Adapter-based Extension Of Multi-speaker Text-to-speech Model For New Speakers (2022)6.77