Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Vision-language foundation model for generalizable nasal disease diagnosis using unlabeled endoscopic records

作者:Xueli Liu, Wentao Gong, Xiao Chen, Xiao Chen, Zhen Li, Yinlong Liu, Li Wang, Quan Liu, Xicai Sun, Xiaofeng Liu, Xinrong Chen, Xinrong Chen, Yuxuan Shi, Hongmeng Yu · 发表于:Pattern Recognition · 年份:2025 · DOI:10.1016/j.patcog.2025.111646 · 被引用次数:11 · 研究领域:Nasal Surgery and Airway Studies、Sinusitis and nasal conditions

Medical artificial intelligence (AI) holds significant potential in identifying signs of health conditions in nasal endoscopic images, thereby accelerating the diagnosis of diseases and systemic disorders. However, the performance of AI models heavily relies on expert annotations, and these models are usually task-specific with limited generalization performance across various clinical applications. In this paper, we introduce NasVLM, a Nasal Vision-Language foundation Model designed to extract universal representations from unlabeled nasal endoscopic data. Additionally, we construct a large-scale nasal endoscopic pre-training dataset and three downstream validation datasets from routine diagnostic records. The core strength of NasVLM lies in its ability to learn cross-modal semantic representations and perform multi-granular report-image alignment without depending on expert annotations. Furthermore, to the best of our knowledge, it is the first medical foundation model that effectively aligns medical report with multiple images of different anatomic regions, facilitated by a well-designed hierarchical report-supervised learning framework. The experimental results demonstrate that NasVLM has superior generalization performance across diverse diagnostic tasks and surpasses state-of-the-art self- and report-supervised methods in disease classification and lesion localization, especially in scenarios requiring label-efficient fine-tuning. For instance, NasVLM can distinguish no...