Robust Audio-visual Target Speaker Extraction With Emotion-aware Multiple Enrollment Fusion
2025 Β· Zhan Jin, Bang Zeng, Peijun Yang, et al.
Abstract
Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images, and frame-level lip motion or facial expression features --can significantly improve performance. However, real-world applications often suffer from intermittent signal loss, especially for frame-level cues. This paper systematically investigates the robustness of multi-enrollment fusion under varying degrees of modality missing. Results show that while full multimodal fusion excels under ideal conditions, its performance degrades sharply when encountering unseen modalities missing during the testing. Crucially, training with a high missing rate dramatically enhances robustness, maintaining stable performance even under severe test-time modality missing. We demonstrate that fusing the complementary one frame of face image with frame-level lip features achieves both strong performance and robustness for the A
Authors
(none)
Tags
Stats
Related papers
- Multi-layer Feature Fusion Convolution Network For Audio-visual Speech Enhancement (2021)0.00
- Attentive Fusion Enhanced Audio-visual Encoding For Transformer Based Robust Speech Recognition (2020)0.00
- Enhancing Real-world Active Speaker Detection With Multi-modal Extraction Pre-training (2024)5.24
- Audio-guided Fusion Techniques For Multimodal Emotion Analysis (2024)4.52
- Audio-visual Target Speaker Enhancement On Multi-talker Environment Using Event-driven Cameras (2019)8.09
- Active Speaker Detection As A Multi-objective Optimization With Uncertainty-based Multimodal Fusion (2021)7.50
- Target Speech Extraction With Pre-trained Av-hubert And Mask-and-recover Strategy (2024)4.52
- Av-sepformer: Cross-attention Sepformer For Audio-visual Target Speaker Extraction (2023)0.00