Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Automating the optimization of proton PBS treatment planning for head and neck cancers using policy gradient‐based deep reinforcement learning

作者:Qingqing Wang, Chang Chang · 发表于:Medical Physics · 年份:2025 · DOI:10.1002/mp.17654 · 被引用次数:7 · 研究领域:Radiation Therapy and Dosimetry、Advanced Radiotherapy Techniques、Prostate Cancer Diagnosis and Treatment

BACKGROUND: Proton pencil beam scanning (PBS) treatment planning for head and neck (H&N) cancers is a time-consuming and experience-demanding task where a large number of potentially conflicting planning objectives are involved. Deep reinforcement learning (DRL) has recently been introduced to the planning processes of intensity-modulated radiation therapy (IMRT) and brachytherapy for prostate, lung, and cervical cancers. However, existing DRL planning models are built upon the Q-learning framework and rely on weighted linear combinations of clinical metrics for reward calculation. These approaches suffer from poor scalability and flexibility, that is, they are only capable of adjusting a limited number of planning objectives in discrete action spaces and therefore fail to generalize to more complex planning problems. PURPOSE: Here we propose an automatic treatment planning model using the proximal policy optimization (PPO) algorithm in the policy gradient framework of DRL and a dose distribution-based reward function for proton PBS treatment planning of H&N cancers. METHODS: The planning process is formulated as an optimization problem. A set of empirical rules is used to create auxiliary planning structures from target volumes and organs-at-risk (OARs), along with their associated planning objectives. Special attention is given to overlapping structures with potentially conflicting objectives. These planning objectives are fed into an in-house optimization engine to generat...