Scholay

学术搜索 · AI 审稿 · LaTeX 协作

AutoPET Challenge on Fully Automated Lesion Segmentation in Oncologic PET/CT Imaging, Part 2: Domain Generalization

作者:Jakob Dexl, Sergios Gatidis, Marcel Früh, Katharina Jeblick, Andreas Mittermeier, Anna Theresa Stüber, Balthasar Schachtner, Johanna Topalis, Matthias P. Fabritius, Sijing Gu, Gowtham Krishnan Murugesan, Jeff VanOss, Jin Ye, Junjun He, Anissa Alloula, Bartłomiej W. Papież, Zacharia Mesbah, Romain Modzelewski, Matthias Hadlich, Zdravko Marinov, Rainer Stiefelhagen, Fabian Isensee, Klaus H. Maier-Hein, Adrián Galdrán, Konstantin Nikolaou, Christian la Fougère, Moon Young Kim, Nico Kallenberg, Jens Kleesiek, Ken Herrmann, Rudolf A. Werner, Michael Ingrisch, Clemens C. Cyran, Thomas Küstner · 发表于:Journal of Nuclear Medicine · 年份:2025 · DOI:10.2967/jnumed.125.270260 · 被引用次数:6 · 研究领域:Radiomics and Machine Learning in Medical Imaging、Medical Imaging Techniques and Applications、AI in cancer detection

This article reports the results of the second iteration of the autoPET challenge on automated lesion segmentation in whole-body PET/CT, held in conjunction with the 26th International Conference on Medical Image Computing and Computer Assisted Intervention in 2023. In contrast to the first autoPET challenge, which served as a proof of concept, this study investigates whether machine learning–based segmentation models trained on data from a single source can maintain performance across clinically relevant variations in PET/CT data, reflecting the demands of real-world deployment. Methods: A comprehensive biomedical segmentation challenge on PET/CT domain generalization was designed and conducted. Participants were tasked to train machine learning models on annotated whole-body 18 F-FDG data ( n = 1,014). These models were then evaluated on a test set of 200 samples from 5 clinically relevant domains, including variations in institutions, pathologies, and populations and a different tracer. Performance was measured in terms of average dice similarity coefficient, average false-positive volume, and average false-negative volume. The best-performing teams were awarded in 3 categories. Furthermore, a detailed analysis was conducted after the challenge, examining results across domains and unique instances, along with a ranking analysis. Results: Generalization from a single-source domain remains a significant challenge. Seventeen international teams successfully participate...