← all papers · overview

CLASP: Cross-modal Salient Anchor-based Semantic Propagation For Weakly-supervised Dense Audio-visual Event Localization

Abstract

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting \textit\{cross-modal salient anchors\}, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a \textit\{Mutual Event Agreement Evaluation\} module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a \textit\{Cross-modal Salient Anchor Identification\} module, which identifies the audio and visual anchor features through global-vide

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).