Scholay

学术搜索 · AI 审稿 · LaTeX 协作

MambaSTF: Visual State Space Model for Remote Sensing Images Spatiotemporal Fusion

作者:Si-Chen Lu, Juanjuan Jing, Kan Wei, Lei Yang, Boyang Nie, Ming-Fei Li, Jin-Song Zhou · 发表于:IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · 年份:2026 · DOI:10.1109/jstars.2026.3668857

Spatiotemporal fusion (STF) methods are pivotal for generating satellite imagery with high spatial and temporal resolution, enabling continuous, gap-free monitoring of land surface dynamics. While deep learning-based STF has shown remarkable potential, it often struggles to balance computational efficiency with effective long-range dependency modeling and adaptive feature extraction across heterogeneous landscapes. To address these issues, we introduce MambaSTF, a novel STF framework based on visual state space models, capable of capturing long-range contexts with linear complexity. Its core includes a Quad-directional 2D-Selective-Scan (QD-SS2D) mechanism that systematically models anisotropic spatial structures via multidirectional scanning. The dedicated State Space Fusion (SSF) module employs dual orthogonal streams: Spatial Selective-Scan Fusion (S-SSF) interleaves multiresolution features to enhance context, while Temporal Selective-Scan Fusion (T-SSF) aligns multitemporal sequences to capture dynamics. These are adaptively fused through a gating mechanism, enabling context-aware spatiotemporal integration. Extensive experiments across diverse datasets confirm its superiority in reconstructing complex landscapes under varying atmospheric conditions. Rigorous uncertainty quantification further demonstrates operational stability and robustness, even amid abrupt land cover changes.