Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Boosting Multimodal Remote Sensing Image Classification With Transformer-Based Heterogeneously Salient Graph Representation

作者:Jiaqi Yang, Bo Du, Rong Liu, Zhu Mao, L Zhang · 发表于:IEEE Transactions on Geoscience and Remote Sensing · 年份:2026 · DOI:10.1109/tgrs.2026.3686762 · 被引用次数:6 · 研究领域:Visual Attention and Saliency Detection、Remote-Sensing Image Classification、Automated Road and Building Extraction

Data collected by different modalities can provide a wealth of complementary information, such as hyperspectral image (HSI) to offer rich spectral-spatial properties, synthetic aperture radar (SAR) to provide structural information about the Earth's surface, and light detection and ranging (LiDAR) to cover altitude information about ground elevation. Therefore, a natural idea is to combine multimodal images for refined and accurate land-cover interpretation. Although many efforts have been attempted to achieve multi-source remote sensing image classification, there are still three issues as follows: 1) indiscriminate feature representation without sufficiently considering modal heterogeneity, 2) abundant features and complex computations associated with modeling long-range dependencies, and 3) overfitting phenomenon caused by sparsely labeled samples. To overcome the above barriers, a transformer-based heterogeneously salient graph representation (THSGR) approach is proposed in this paper. First, a multimodal heterogeneous graph encoder is presented to encode distinctively non-Euclidean structural features from heterogeneous data. Then, a self-attention-free multi-convolutional modulator is designed for effective and efficient long-term dependency modeling. Finally, a mean forward strategy is developed in order to avoid overfitting. Based on the above structures, the proposed model is able to break through modal gaps to obtain differentiated graph representation with competit...