Scholay

学术搜索 · AI 审稿 · LaTeX 协作

THMM-CLIP: Task-Guided Hierarchical Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning

作者:Yingdong Pan, Zhaoquan Yuan, Xiao Wu, Zechao Li, Changsheng Xu · 发表于:ACM Transactions on Multimedia Computing Communications and Applications · 年份:2025 · DOI:10.1145/3785477 · 被引用次数:1 · 研究领域:Domain Adaptation and Few-Shot Learning、Multimodal Machine Learning Applications、Face recognition and analysis

Class incremental learning (CIL) requires models to acquire knowledge from sequential tasks containing non-overlapping classes while avoiding catastrophic forgetting. While vision-language foundation models like CLIP demonstrate remarkable potential for CIL through their pre-trained cross-modal alignment capabilities, existing CLIP-based approaches critically overlook the progressive degradation of visual representations in incremental scenarios . Through feature space analysis, we identify a crucial dichotomy : textual embeddings maintain stable discriminative power across sequential tasks, whereas visual features exhibit progressive deterioration manifested by intra-task confusion (ambiguous decision boundaries between co-occurring classes) and inter-task interference (semantic collision between historical and novel categories). To address these dual challenges, we propose task-guided hierarchical multi-modal alignment (THMM-CLIP), a framework that establishes persistent visual-textual coherence through hierarchical multi-modal alignment (HMA) and robust prompt selection (RPS). HMA adapts lightweight task-specific prompt vectors to dynamically recalibrate the CLIP image encoder, thereby achieving: (i) intra-task alignment, (ii) inter-task discriminability alignment, and (iii) global structural alignment with textual features. RPS incorporates a dual-level task identifier that integrates class-level and task-level representative features to ensure precise prompt retrieval du...