Scholay

学术搜索 · AI 审稿 · LaTeX 协作

The Power of Modality: Improving Polyp Segmentation With Multimodal Information

作者:Fang Wang, Pu Wang, Meng Zhao, Chenggang Shan, Zhen Yang · 发表于:IET Image Processing · 年份:2026 · DOI:10.1049/ipr2.70305 · 被引用次数:1 · 研究领域:Generative Adversarial Networks and Image Synthesis、Colorectal Cancer Screening and Detection、Multimodal Machine Learning Applications

ABSTRACT Accurate polyp segmentation from a single color image remains a significant challenge due to the complex appearance of lesions and the lack of diverse contextual priors. Existing methods usually rely on a limited image prior, which leads to suboptimal results. We propose a novel approach that leverages rich contextual information from multiple modalities (including depth, higher order semantics, and edge features) to produce precise segmentation results. By integrating the Depth Anything model, Large Language Models, and Edge Extractor, our method effectively fuses these diverse priors to overcome the limitations of single modality approaches. In addition, we introduce a diffusion modeling framework to bring powerful generative priors for endoscopic image segmentation. This is a flexible deep network architecture that efficiently fuses multimodal information and can accommodate any number of multimodal inputs. Extensive experimental results demonstrate that our approach achieves promising performance on several benchmarks.