Distribution-aware reincarnating reinforcement learning mitigates dual distribution shifts in teacher-guided offline-to-online learning
作者:Jiahao Xue, Zhuxiao Wang, Hong Wang, Ying Zhang, Yun Ju · 发表于:Scientific Reports · 年份:2026 · DOI:10.1038/s41598-026-63862-9 · 研究领域:Reinforcement Learning in Robotics、Neural Networks and Reservoir Computing、Advanced Bandit Algorithms Research
Offline pretraining followed by online fine-tuning is a classical paradigm in reinforcement learning. When a pretrained teacher is available, this paradigm can be extended from generic initialization to efficient teacher-guided learning. Reincarnating Reinforcement Learning (RRL) is a representative framework in this setting. However, in teacher-guided offline-to-online learning, the transferred teacher knowledge may become unreliable under distribution shift. In particular, the offline stage suffers from out-of-distribution (OOD) action extrapolation on teacher replay, while the online stage suffers from replay-distribution shift during the transition from teacher-generated to student-generated experience. We address these issues with Distribution-Aware Reincarnating Reinforcement Learning, which combines teacher-guided conservative learning for reliable offline value reuse and an adaptive balanced replay buffer for stable online replay transition. Experiments on six Atari tasks show that the proposed method achieves strong overall performance, improves Q-value conservatism, and yields smoother offline-to-online transfer.