Learning To Unify Audio, Visual And Text For Audio-enhanced Multilingual Visual Answer Localization
2024 Β· Zhibin Wen, Bin Li
Abstract
The goal of Multilingual Visual Answer Localization (MVAL) is to locate a video segment that answers a given multilingual question. Existing methods either focus solely on visual modality or integrate visual and subtitle modalities. However, these methods neglect the audio modality in videos, consequently leading to incomplete input information and poor performance in the MVAL task. In this paper, we propose a unified Audio-Visual-Textual Span Localization (AVTSL) method that incorporates audio modality to augment both visual and textual representations for the MVAL task. Specifically, we integrate features from three modalities and develop three predictors, each tailored to the unique contributions of the fused modalities: an audio-visual predictor, a visual predictor, and a textual predictor. Each predictor generates predictions based on its respective modality. To maintain consistency across the predicted results, we introduce an Audio-Visual-Textual Consistency module. This module
Authors
(none)
Tags
Stats
Related papers
- Sviqa: A Unified Speech-vision Multimodal Model For Textless Visual Question Answering (2025)0.00
- Alignvsr: Audio-visual Cross-modal Alignment For Visual Speech Recognition (2024)0.00
- Unified Video-language Pre-training With Synchronized Audio (2024)0.00
- Cross-modal Global Interaction And Local Alignment For Audio-visual Speech Recognition (2023)7.50
- AV2AV: Direct Audio-visual Speech To Audio-visual Speech Translation With Unified Audio-visual Speech Representation (2023)6.77
- MLCA-AVSR: Multi-layer Cross Attention Fusion Based Audio-visual Speech Recognition (2024)10.07
- Dual Mean-teacher: An Unbiased Semi-supervised Framework For Audio-visual Source Localization (2024)5.24
- Large Language Models Are Strong Audio-visual Speech Recognition Learners (2024)9.59