Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

作者:Leonardo Chimirri, J. Harry Caufield, Yasemin Bridges, Nicolas Matentzoglu, Michael Gargano, Mario Cazalla, Shihan Chen, Daniel Daniš, Alexander J.M. Dingemans, Klara Gehle, Petra Gehle, Adam S L Graefe, Weihong Gu, Markus S. Ladewig, Pablo Lapunzina, Julián Nevado, Enock Niyonkuru, Soichi Ogishima, Dominik Seelow, Jair Tenorio, Marek Turnovec, Bert B.A. de Vries, Kai Wang, Kyran Wissink, Zafer Yüksel, Gabriele Zucca, Melissa Haendel, Chris Mungall, Justin Reese, Peter N. Robinson · 发表于:EBioMedicine · 年份:2025 · DOI:10.1016/j.ebiom.2025.105957 · 被引用次数:10 · 研究领域:Genomics and Rare Diseases、Biomedical Text Mining and Ontologies、Topic Modeling

BACKGROUND: Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. METHODS: We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtype...