Deep modular skip-attention networks for multimodal named entity recognition
作者:Chengen Lai, Shengli Song, Sitong Yan, Guangneng Hu · 发表于:Journal of Information and Intelligence · 年份:2026 · DOI:10.1016/j.jiixd.2026.04.004 · 研究领域:Topic Modeling、Natural Language Processing Techniques、Authorship Attribution and Profiling
Multimodal named entity recognition(MNER) has achieved significant progress in recent years. However, existing approaches still suffer from two drawbacks. Firstly, for interaction effectiveness, the shallow multimodal interaction models fail in capturing fine-grained interactions between multimodal instances and suffer from limited modality information. The deep models that use simple concatenation strategies demand a huge amount of computation resources while achieving better performance. Secondly, for interaction strategies, if the associated image is noisy and unrelated to the text, multimodal representation in MNER is often biased and even omits the important information from the dominant modality in current fusion methods, which makes them suffer from visual bias. To address the above issues, we proposed DESER ( DE ep Modular S kip Attention Networks for MN ER ) to effectively model the interactions across modalities and debias the noisy modality impact. Unlike shallow crossmodal attention models that capture only single-layer interactions, and deep concatenation-based frameworks that indiscriminately fuse modalities at all layers, DESER introduces a modular skip-attention architecture that enables deep, text-guided multimodal interaction while explicitly preserving dominant textual semantics. The proposed Skip-Attention layer creates interlayer shortcuts that selectively bypass visual adaptation, allowing the model to balance interaction depth and computational efficien...