Reliability of uncertainty quantification methods for deep learning auto-segmentation in head and neck organs at risk
作者:Joëlle E. van Aalst, Federica Carmen Maruccio, Rita Simões, Tomas Janssen, Jelmer M. Wolterink, Peter M. A. van Ooijen, Charlotte L. Brouwer · 发表于:Physics in Medicine and Biology · 年份:2025 · DOI:10.1088/1361-6560/ae110c · 被引用次数:5 · 研究领域:Advanced Radiotherapy Techniques、Radiomics and Machine Learning in Medical Imaging、Radiation Therapy and Dosimetry
Abstract Objective. Deep learning auto-segmentation has greatly advanced contouring in radiotherapy. However, quality assurance remains necessary due to performance fluctuation among individual patients. This manual process reintroduces variability and partially reduces time-saving benefits. As a solution, uncertainty quantification (UQ) is increasingly explored for its ability to estimate output confidence. While numerous methods to quantify uncertainty exist, their comparative reliability remains underexplored. This study compares the reliability of commonly used UQ approaches for auto-segmentation in radiotherapy. Approach. We evaluated the reliability of three popular uncertainty methods (Monte Carlo dropout, deep ensemble modelling and test-time augmentation) and uncertainty metrics (predictive entropy, mutual information and variance). We trained a 3D U-Net within the nnU-Net framework to segment 19 organs at risk (OAR) for head and neck cancer patients. We evaluated the reliability of the UQ methods and metrics on a set of 10 patients using segmentation model accuracy (surface Dice similarity coefficient), confidence calibration (expected calibration error (label)), and error localisation ability (uncertainty-error (U-E) overlap). Both multi-class and class-specific uncertainty maps were assessed. Main results. Segmentation accuracy remained stable without significant deviations across all UQ methods with respect to the baseline model without UQ. The reliability of dif...