Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Multi-modal cross Swin transformer network for multi-label classification landslide detection with optical and SAR images of Luding

作者:Yongxin Li, Yukun Xue, Zhihui Xin, Guisheng Liao, Penghui Huang · 发表于:International Journal of Applied Earth Observation and Geoinformation · 年份:2025 · DOI:10.1016/j.jag.2025.104954 · 被引用次数:2 · 研究领域:Landslides and related hazards、Remote-Sensing Image Classification、Synthetic Aperture Radar (SAR) Applications and Techniques

Natural phenomena such as earthquakes and heavy rainfall can trigger landslide events hundreds or even thousands of times in a given area. Therefore, rapid and efficient intervention in affected regions is essential. Existing studies have demonstrated good detection performance using optical remote sensing images; however, optical data has significant limitations under cloud cover. Moreover, current multi-modal fusion methods struggle to effectively capture the nonlinear interactions between different modes when processing data with large informational disparities. To address these issues, we developed the China Luding multi-modal landslide dataset, which includes optical and polarimetric synthetic aperture radar (PolSAR) images. We also propose a multi-modal network for landslide detection based on multi-label classification, called multi-modal cross Swin Transformer network (MCSTNet). Additionally, a new weighted asymmetric loss function (WASL) is proposed to improve multi-label classification tasks. Our proposed network consists of two stages: feature extraction and feature fusion. In the first stage, two independent branches extract high-level semantic features from optical and PolSAR images. In the second stage, a multi-modal cross multi-head self-attention (MCMSA) mechanism fuses the high-level semantic features from the multi-modal information. Therefore, the model enhances the learning and feature representation capabilities through parallel processing the input data ...