SegORG : Report Generation of Oral Potentially Malignant Disorders Image Based on Lesion Segmentation‐Enhanced LLM
作者:Rui Zhang, Peng Huang, Tingting Ding, Yaowu Chen, Xiang Tian, Yuqi Cao, Wei Chen, Xiaoyan Chen, Qianming Chen, Fudong Zhu · 发表于:Oral Diseases · 年份:2025 · DOI:10.1111/odi.70148 · 被引用次数:2 · 研究领域:AI in cancer detection、Radiomics and Machine Learning in Medical Imaging、Head and Neck Cancer Studies
OBJECTIVES: To develop an automated system for generating standardized reports for oral potentially malignant disorders (OPMDs) from white-light images, aiming to reduce documentation workload, facilitate early intervention, and enable longitudinal lesion monitoring. METHODS: We proposed the SegORG model using 441 oral mucosa images and 1323 corresponding reports. It employed SegFormer for lesion segmentation; a visual encoder extracted global and local visual embeddings, which were projected into a pre-trained large language model (LLM feature) space via a lightweight visual mapper. The Qwen2.5-7B model then generated structured diagnostic reports, enhanced by text augmentation techniques to improve diversity and professionalism. RESULTS: SegORG achieved BLEU-4, ROUGE-L, and CIDEr scores of 0.291, 0.517, and 0.578, respectively. Additionally, the model obtained a clinical diagnostic F1-score of 0.695 and a median expert rating of 4 on the Likert scale (p < 0.001). It significantly outperformed conventional baseline models (R2Gen, METransformer, and SwinB+BERT9k) and contemporary general-purpose multimodal LLMs (GPT-4, Gemini 2.5 Pro, and Qwen2.5-VL), despite its lean architecture (90.4 M trainable parameters). CONCLUSIONS: By enhancing visual feature extraction and achieving efficient text alignment, SegORG offers an effective technical pathway for OPMDs reports automation. While single-center validation shows promise, multicenter trials are needed to assess generalizability...