Conservative Offline Meta-Reinforcement Learning with Task Similarity Measurement
作者:Haorui Li, Jiaqi Liang, Linjing Li, Daniel Zeng · 年份:2025 · DOI:10.1109/icassp49660.2025.10888196 · 被引用次数:1 · 研究领域:Reinforcement Learning in Robotics
Offline meta-reinforcement learning (OMRL) enables reinforcement learning (RL) agents to adapt to unseen tasks without interacting with the environment. However, OMRL faces challenges such as Q-function overestimation and difficulties in inferring tasks correctly and robustly due to distribution discrepancy. In this paper, we introduce ConseRvative q-learning and task similarity mEAsuremenT for Offline meta-Reinforcement learning (CREATOR), a method to address these challenges using only offline datasets, without requiring additional interactions. To mitigate Q-function overestimation, we incorporate conservative Q-learning during training. We also propose a novel task similarity-based distance metric to improve the robustness of task inference. Experimental results demonstrate that the proposed CREATOR effectively reduces Q-function estimation errors, enhances task inference accuracy, and improves generalization performance across a range of challenging domains compared to existing methods.