Dual-Prompt Learning with Cross-Modal Decoders for Few-Shot Whole Slide Image Classification
作者:Yiming Xu, Bingbing Zhang, Wen Zhu, Lizhi Zhang, Bin Liu, Jianxin Zhang, Qiang Zhang · 发表于:IEEE International Conference on Bioinformatics and Biomedicine · 年份:2025 · DOI:10.1109/bibm66473.2025.11356780 · 研究领域:Computer Science
Few-shot learning offers a promising solution for computational pathology by alleviating the reliance on large an-notated datasets, but faces challenges from the high redundancy in whole slide images and underutilized cross-modal knowledge. Existing methods typically use foundation models only for pre-liminary feature extraction while employing fixed or single-level prompts that lack multi-scale pathological representation. To address these limitations, we propose a Hierarchical Vision-Text Prompt (H- VTP) framework that enables multi-level cross-modal interaction through GPT-4 generated Local Instance Prompts for patch-level morphological details and Global Semantic Prompts for slide-level diagnostic context. A dual-branch decoding mech-anism with Text-Guided-Patch Decoder and Patch-Augmented-Text Decoder facilitates closed-loop vision-text fusion, while a parameter-efficient adaptation strategy trains only lightweight prompts and adapters. Extensive experiments on three cancer subtype datasets demonstrate the superiority of H-VTP in few-shot WSI classification, confirming its effectiveness for clinical applications.