Deep Audio-visual Singing Voice Transcription Based On Self-supervised Learning Models
2023 Β· Xiangming Gu, Wei Zeng, Jianan Zhang, et al.
Abstract
Singing voice transcription converts recorded singing audio to musical notation. Sound contamination (such as accompaniment) and lack of annotated data make singing voice transcription an extremely difficult task. We take two approaches to tackle the above challenges: 1) introducing multimodal learning for singing voice transcription together with a new multimodal singing dataset, N20EMv2, enhancing noise robustness by utilizing video information (lip movements to predict the onset/offset of notes), and 2) adapting self-supervised learning models from the speech domain to the singing voice transcription task, significantly reducing annotated data requirements while preserving pretrained features. We build a self-supervised learning based audio-only singing voice transcription system, which not only outperforms current state-of-the-art technologies as a strong baseline, but also generalizes well to out-of-domain singing data. We then develop a self-supervised learning based video-only s
Authors
(none)
Tags
Stats
Related papers
- Semi-supervised Learning For Singing Synthesis Timbre (2020)3.58
- Visinger2+: End-to-end Singing Voice Synthesis Augmented By Self-supervised Learning Representation (2024)4.52
- A Melody-unsupervision Model For Singing Voice Synthesis (2021)5.84
- Toward Leveraging Pre-trained Self-supervised Frontends For Automatic Singing Voice Understanding Tasks: Three Case Studies (2023)0.00
- Automatic Lyrics Transcription Using Dilated Convolutional Neural Networks With Self-attention (2020)10.07
- Learn To Sing By Listening: Building Controllable Virtual Singer By Unsupervised Learning From Voice Recordings (2023)0.00
- Unsupervised Singing Voice Conversion (2019)11.19
- Self-supervised Singing Voice Pre-training Towards Speech-to-singing Conversion (2024)0.00