FWDNNet: Cross-Heterogeneous Encoder Fusion via Feature-Level TensorDot Operations for Land-Cover Mapping
作者:Boaz Mwubahimana, Yan Jianguo, Dingruibo Miao, Swalpa Kumar Roy, Zhuohong Li, Le Ma, Clarisse Kagoyire, Haonan Guo, Maurice Mugabowindekwe, E Nyandwi, Isaac Nzayisenga, Hafashimana Athanase, Eugene Maridadi, Jean Baptiste Nsengiyumva, Elie Byukusenge, Remy Dukundane, Gaspard Rwanyiziri, Xiao Huang · 发表于:IEEE Transactions on Geoscience and Remote Sensing · 年份:2026 · DOI:10.1109/tgrs.2026.3652451 · 被引用次数:2 · 研究领域:Remote-Sensing Image Classification、Remote Sensing in Agriculture、Advanced Image Fusion Techniques
Autonomous feature extraction from high-resolution (HR) Remote Sensing (RS) imagery presents complex spatial–spectral heterogeneity and diverse land-cover patterns, challenging conventional single-architecture models that lack multiscale feature diversity and cross-domain generalization. Convolutional neural networks (CNNs) effectively capture localized hierarchical structures, whereas transformers model long-range contextual dependencies via self-attention. However, using these architectures independently restricts their complementary potential. This study proposes FWDNNet (Fused Weights Deep Neural Tokenization Networks), a fusion framework introducing a CNN-to-token conversion mechanism that transforms convolutional feature maps from heterogeneous encoders into transformer-compatible token sequences. Beyond standard concatenation, FWDNNet integrates a TensorDot-based fusion module to learn high-order cross-architectural interactions, while a shared decoding pathway maintains multiscale feature hierarchies and computational efficiency. Experiments across three diverse datasets show that FWDNNet achieves 58.2 ms inference latency with minimal cross-domain performance degradation (2.3–4.2% mIoU drop). Compared to conventional fusion strategies, FWDNNet attains 21% faster inference and 1.7–2.7% higher mIoU over baseline methods. See Materials at https://github.com/BoazGithub/FWDNNet for reproducibility.