Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Locality-Aware Cross-Modal Correspondence Learning for Dense Audio-Visual Events Detection

作者:Ling Xing, Hongyu Qu, Rui Yan, Xiangbo Shu, Jinhui Tang · 发表于:IEEE Transactions on Circuits and Systems for Video Technology · 年份:2025 · DOI:10.1109/tcsvt.2025.3629609 · 被引用次数:9 · 研究领域:Multimodal Machine Learning Applications、Speech and Audio Processing、Music and Audio Processing

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However, complex audio-visual scenes often involve asynchronization between modalities, making accurate localization challenging. Existing DAVE solutions extract audio and visual features through unimodal encoders, and fuse them via dense cross-modal interaction. However, independent unimodal encodingstruggles to emphasize shared semantics between modalitieswithout cross-modal guidance, while dense cross-modal attention mayover-attend to semantically unrelated audio-visual features. To address these problems, we present LOCO, a Locality-aware cross-modal Correspondence learning framework for DAVE. LOCO leverages the local temporal continuity of audio-visual events as important guidance to filter irrelevant cross-modal signals and enhance cross-modal alignment throughout both unimodal and cross-modal encoding stages. i) Specifically, LOCO applies Local Correspondence Feature (LCF) Modulation to enforce unimodal encoders to focus on modality-shared semantics by modulating agreement between audio and visual features based on local cross-modal coherence. ii) To better aggregate cross-modal relevant features, we further customize Local Adaptive Cross-modal (LAC) Interaction, which dynamically adjusts attention regions in a data-driven manner. This adaptive m...