Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Label-Semantic-Based Prompt Tuning for Vision Transformer Adaptation in Medical Image Analysis

作者:Yu Bai, Liang Bai, Xian Yang, Jiye Liang · 发表于:IEEE Transactions on Circuits and Systems for Video Technology · 年份:2025 · DOI:10.1109/tcsvt.2025.3578111 · 被引用次数:4 · 研究领域:Image Retrieval and Classification Techniques、AI in cancer detection、Medical Image Segmentation Techniques

Adapting Vision Transformers (ViTs) to medical image analysis is challenging due to the scarcity of annotated data and the significant domain shift from natural to medical images. Traditional fine-tuning approaches, while effective, require storing separate model parameters for each task, leading to high computational costs. Existing prompt tuning methods reduce this overhead by introducing task-specific prompt tokens, but they often fail to fully leverage label semantics, resulting in suboptimal performance for medical tasks. To address these limitations, we propose a label-semantic-based prompt tuning method (LPT), which transforms the visual prompt learning problem into a text-image alignment task. Unlike traditional prompt methods that only focus on visual prompts, LPT incorporates label semantics through a cross-attention-based module to better align image features with the target labels. This approach not only captures rich semantic information from the labels but also enhances the model’s ability to extract fine-grained image details relevant to specific medical conditions. By leveraging label-text alignment during training, LPT improves both label utilization and model adaptability, enabling more accurate predictions. Extensive experiments on eight diverse medical datasets demonstrate that LPT significantly improves diagnostic accuracy and generalization, outperforming both traditional fine-tuning and current prompt-based methods, especially in data-limited scenarios.