Alignclip: navigating the misalignments for robust vision-language generalization
作者:Zhongyi Han, Gongxu Luo, Hao Sun, Yaqian Li, Bo Han, Mingming Gong, Kun Zhang, Tongliang Liu · 发表于:Machine Learning · 年份:2025 · DOI:10.1007/s10994-025-06742-z · 被引用次数:7 · 研究领域:Multimodal Machine Learning Applications、Domain Adaptation and Few-Shot Learning、Advanced Image and Video Retrieval Techniques
In the realm of Vision-Language Pretraining models, achieving robust and adaptive representations is a cornerstone for successfully handling the unpredictability of real-world scenarios. This paper delves into two pivotal misalignment challenges inherent to Contrastive Language-Image Pre-training (CLIP) models: attention misalignment, which leads to an overemphasis on background elements rather than salient objects, and predictive category misalignment, characterized by the model’s struggle to discern between classes based on similarity. These misalignments undermine the representational stability essential for dynamic, real-world applications. To address these challenges, we propose AlignCLIP, an advanced fine-tuning methodology distinguished by its attention alignment loss, designed to calibrate the distribution of attention across multi-head attention layers. Furthermore, AlignCLIP introduces semantic label smoothing, a technique that leverages textual class similarities to refine prediction hierarchies. Through comprehensive experimentation on a variety of datasets and in scenarios involving distribution shifts and unseen classes, we demonstrate that AlignCLIP significantly enhances the stability of representations and shows superior generalization capabilities.