Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Few-Shot Object Detection in Remote Sensing Images via Dynamic Adversarial Contrastive-Driven Semantic-Visual Fusion

作者:Yang Xu, Jiahui Qin, Tianming Zhan, Huapeng Wu, Zhihui Wei, Zebin Wu · 发表于:IEEE Transactions on Geoscience and Remote Sensing · 年份:2025 · DOI:10.1109/tgrs.2025.3602969 · 被引用次数:4 · 研究领域:Image Processing Techniques and Applications、Infrared Target Detection Methodologies、Advanced Image Processing Techniques

The acquisition of remote sensing images (RSIs) requires expensive equipment, such as satellites or aircraft, along with advanced sensing technologies. Additionally, the complexity of terrains and the inaccessibility of certain regions further exacerbate the difficulty of obtaining remote sensing image data, particularly for novel classes. Therefore, this paper focuses on two core challenges in few-shot object detection (FSOD) for RSIs: (1) The difficulty in feature learning caused by insufficient samples of novel classes; (2) Large intra-class variability and small inter-class differences in novel classes, referring to the feature similarity between different categories and the feature variability within the same category. To address the above two major challenges, we propose a FSOD framework in RSIs that integrates the Semantic-Visual Fusion Module (SVFM) with the Dynamic Adversarial Contrastive Loss Function (DACL). Firstly, to tackle the issue of insufficient feature representation for novel classes, we introduce a text inference branch that utilizes the CLIP text encoder to generate rich semantic representations. Furthermore, to promote deep integration and synergistic effects between textual information and image features, we design a Multi-Level Cross Attention Mechanism (MLCA) and a Dual-Head Co-Attention Guidance Feature Fusion Module (DHCA) to enhance the semantic understanding and visual representation capabilities for novel classes, effectively compensating for th...