M-CGL: Mamba-Enhanced Concept Guided Learning for fine-Grained Image Classification
作者:Xiang Xue, Yatu Ji, 仁庆道尔吉, Bao Shi, Min Lu, Yuheng Guo, Nier Wu, Ganqiqige Cha · 年份:2026 · DOI:10.1109/icassp55912.2026.11464388 · 研究领域:Image Retrieval and Classification Techniques、Machine Learning and Data Classification、Machine Learning and Algorithms
The key challenge in Fine-Grained Visual Classification (FGVC) lies in capturing local subtle differences. Although Transformers excel at modeling long-range dependencies, their patch partitioning scheme in image processing tends to weaken the association between local and global features. While sliding windows alleviate neighborhood fragmentation, they still struggle to avoid the block-like discretization of features. To address this issue, we propose the Mamba Concept-Guided Learning (M-CGL) framework, which consists of two novel components: the Mamba Semantic Concept Modeling (M-SCM) module and the Mamba Semantic Concept Fusion (M-SCF) module. The M-SCM module enhances the inter-relationships among fine-grained features to extract more discriminative representations. The M-SCF module fuses discriminative features and feature maps across multiple stages, enabling hierarchical concept alignment and preserving spatial-semantic consistency throughout the network. While M-SCF ensures intra-sample semantic coherence, it neglects inter-sample structural relations. Thus we adopt SoftTriple loss to explicitly enforce intra-class compactness and inter-class separability, enhancing discrimination among visually similar categories. Experiments show that our method achieves accuracy gains of 0.54% on CUB-200, 1.56% on Aircraft and 5.5% on Fiber, respectively.