Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Integrating Global and Local Information for Remote Sensing Image–Text Retrieval

作者:Ziyun Chen, Fan Liu, Zhan-Rong Guan, Qiang Zhou, Xiaocong Zhou, Chuanyi Zhang · 发表于:IEEE Geoscience and Remote Sensing Letters · 年份:2025 · DOI:10.1109/lgrs.2025.3616154 · 被引用次数:2 · 研究领域:Computer Science

Pretrained vision–language models (VLMs) have demonstrated promising performance in remote sensing (RS) image–text retrieval tasks. However, the scarcity of high-quality image–text datasets remains a challenge in fine-tuning VLMs for RS. The captions in existing datasets tend to be uniform and lack details. To fully use rich detailed information from RS images, we propose a method to fine-tune VLMs. We first construct a new visual–language dataset that balances both global and local information for RS (GLRS) image–text retrieval. Specifically, a multimodal large language model (MLLM) is used to generate captions for local patches and global captions for the entire image. To effectively use local information, we propose a global and local image captioning method (GLCap). With a large language model (LLM), we further obtain higher quality captions by merging both global and local captions. Finally, we fine-tune the weights of RS-M-contrastive language image pretraining (CLIP) with a progressive global–local fine-tuning strategy on GLRS. Experimental results demonstrate that our method outperforms state-of-the-art (SoTA) approaches on two common RS image–text retrieval downstream tasks. Our code and dataset are available at https://github.com/hhu-czy/GLRS