Improved Speech Reconstruction From Silent Video
2017 Β· Ariel Ephrat, Tavi Halperin, Shmuel Peleg
Abstract
Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible and natural-sounding acoustic speech signal from silent video frames of a speaking person. We train our model on speakers from the GRID and TCD-TIMIT datasets, and evaluate the quality and intelligibility of reconstructed speech using common objective measurements. We show that speech predictions from the proposed model attain scores which indicate significantly improved quality over existing models. In addition, we show promising results towards reconstructing speech from an unconstrained dictionary.
Authors
(none)
Tags
Stats
Related papers
- Vid2speech: Speech Reconstruction From Silent Video (2017)14.15
- Video-driven Speech Reconstruction Using Generative Adversarial Networks (2019)11.39
- Let There Be Sound: Reconstructing High Quality Speech From Silent Videos (2023)6.34
- Speech Reconstruction From Silent Tongue And Lip Articulation By Pseudo Target Generation And Domain Adversarial Training (2023)5.84
- Lipvoicer: Generating Speech From Silent Videos Guided By Lip Reading (2023)3.89
- From Faces To Voices: Learning Hierarchical Representations For High-quality Video-to-speech (2025)0.00
- Lipsound2: Self-supervised Pre-training For Lip-to-speech Reconstruction And Lip Reading (2021)11.39
- Audio-visual Speech Codecs: Rethinking Audio-visual Speech Enhancement By Re-synthesis (2022)15.58