Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Domain-Separated Bottleneck Attention Fusion Framework for Multimodal Emotion Recognition

作者:Peng He, Jun Yu, Chengjie Ge, Ye Yu, W. L. Xu, Lei Wang, Tianyu Liu, Zhen Kan · 发表于:ACM Transactions on Multimedia Computing Communications and Applications · 年份:2025 · DOI:10.1145/3711865 · 被引用次数:11 · 研究领域:Emotion and Mood Recognition、Human Pose and Action Recognition、Video Surveillance and Tracking Methods

As a focal point of research in various fields, human body language understanding has long been a subject of intense interest. Within this realm, the exploration of emotion recognition through the analysis of facial expressions, voice patterns, and physiological signals holds significant practical value. Compared with unimodal approaches, multimodal emotion recognition models leverage complementary information from vision, acoustic, and language modalities to robust perceive the human sentiment attitudes. However, the heterogeneity among modality signals leads to significant domain shifts, posing challenges for achieving balanced fusion. In this article, we propose a Domain-Separated Bottleneck Attention (DBA) Fusion Framework for human multimodal emotion recognition with lower computational complexity. Specifically, we partition each modality into two distinct domains: the invariant/private domain. The invariant domain contains crucial shared information, while the private domain aims to capture modality-specific representations. For the decomposed features, we introduce two sets of bottleneck cross-attention modules to effectively utilize the complementarity between domains to reduce redundant information. In each module, we interweave two Fusion Adapter blocks into the Self-Attention Transformer backbone. Each Fusion Adapter block integrates a small group of latent tokens as bridges for inter-modal and inter-domain interactions, mitigating the adverse effects of modality d...