GRID
Emerging5papers using it
2021first seen
The GRID dataset is a benchmark that contains video recordings of speakers articulating a set of predefined sentences, used to evaluate the performance of Lip-to-Speech synthesis systems.
Papers using GRID (5)
- VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via
Speech-Visage Feature SelectionOn the Audio-visual Synchronization for Lip-to-Speech SynthesisSVTS: Scalable Video-to-Speech SynthesisLipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction
and Lip ReadingRobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting
Self-Supervised Representations