Rethinking the Effect of Unimodal Labels in Multimodal Sentiment Analysis
作者:Lingli Zhang, Tianrui Li, Baiyu Lu, Junlin Fang, Desheng Zheng, Wei Zhou, Weide Liu, Fengmao Lv · 发表于:ACM Transactions on Multimedia Computing Communications and Applications · 年份:2026 · DOI:10.1145/3796718 · 被引用次数:3 · 研究领域:Sentiment Analysis and Opinion Mining、Emotion and Mood Recognition、Multimodal Machine Learning Applications
Multimodal sentiment analysis aims to comprehensively understand human sentiment by integrating diverse modalities, such as text, audio, and vision. To improve modality complementarity, the recent Multimodal Multi-task Learning (MML) framework employs joint training of unimodal and multimodal sentiment analysis tasks using sub-annotations of modality. In this work, we further draw attention to the observation that integrating unimodal tasks may introduce conflicting task information, negatively affecting the multimodal task performance. Motivated by this issue, we propose the Multimodal Task Correlation-aware Learning (MTCL) framework to leverage beneficial task correlations and suppress harmful ones. Specifically, MTCL introduces a Correlation-Adaptive Training (CAT) strategy to learn a task-relation aware unimodal encoder for each modality. First, in order to distinguish whether a sample contains conflicting information, CAT strategy incorporates a Dual-Branch Contrast (DBC) module which divides the training set into a beneficial subset and a harmful subset. Based on this division, CAT strategy proposes an adaptive training loss to guide the model in understanding nuanced multitask correlations. The adaptive training loss has two components: (1) For the beneficial subset, a contrastive loss is utilized to improve the model’s ability to extract complementary representations. (2) For the harmful subset, we apply a task-correction loss to mitigate the negative interference cau...