Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Visual Global-Salient-Guided Network for Remote Sensing Image-Text Retrieval

作者:Yangpeng He, Xin Xu, Hongjia Chen, Jinwen Li, Fangling Pu · 发表于:IEEE Transactions on Geoscience and Remote Sensing · 年份:2024 · DOI:10.1109/tgrs.2024.3466389 · 被引用次数:21 · 研究领域:Image Retrieval and Classification Techniques、Advanced Image and Video Retrieval Techniques

Amid the brisk evolution of remote sensing (RS) technology, the domain of RS cross-modal text-image retrieval (RSCTIR) has captivated scholarly interest for its superior adaptability and symbiotic interaction with human operators. However, due to the heterogeneity between image and text data modalities, feature alignment poses a significant challenge. The existing methodologies overlook the sufficient incorporation of structural guidance during the cross-modal feature interaction alignment process to foster alignment between text and image features. In light of this, we propose an innovative approach for RS image-text retrieval task called visual global-salient-guided network (VGSGN), which comprises two branches: the image branch and the text branch. In the image branch, visual global-salient information sensing module (VGSM) is devised to extract visual global and salient features, aiming to enhance the perception capability for complex backgrounds and scenes in RS images. In the text branch, the textual graph enhancement module (TGEM) is crafted to filter out redundant information in the text features and capture the interactions between words within the text. The design of the multiple visual-guided dynamic fusion (MVGF) module aims to leverage the global and salient features of image to guide the text feature, facilitating cross-modal alignment of text and image features. The experimental results on the widely recognized RSICD and RSITMD datasets corroborate the effectiv...