Towards Specific Domain Prompt Learning via Improved Text Label Optimization
作者:Liangchen Liu, Nannan Wang, Decheng Liu, Xi Yang, Xinbo Gao, Tongliang Liu · 发表于:IEEE Transactions on Multimedia · 年份:2024 · DOI:10.1109/tmm.2024.3413318 · 被引用次数:4 · 研究领域:Topic Modeling、Natural Language Processing Techniques、Text and Document Classification Technologies
Prompt learning has emerged as a thriving parameter-efficient fine-tuning technique for adapting pre-trained vision-language models (VLMs) to various downstream tasks. However, existing prompt learning approaches still exhibit limited capability for adapting foundational VLMs to specific domains that require specialized and expert-level knowledge. Since this kind of specific knowledge is primarily embedded in the pre-defined text labels, we infer that foundational VLMs cannot directly interpret semantic meaningful information from these specific text labels, which causes the above limitation. From this perspective, this paper additionally models text labels with learnable tokens and casts this operation into traditional prompt learning framework. By optimizing label tokens, semantic meaningful text labels are automatically learned for each class. Nevertheless, directly optimizing text label still remains two critical problems, i.e., insufficient optimization and biased optimization. We further address these problems by proposing Modality Interaction Text Label Optimization (MITLOp) and Color-based Consistency Augmentation (CCAug) respectively, thereby effectively improving the quality of the optimized text labels. Extensive experiments indicate that our proposed method achieves significant improvements in VLM adaptation on specific domains.