Scholay

学术搜索 · AI 审稿 · LaTeX 协作

NADPEx: An on-policy temporally consistent exploration method for deep\n reinforcement learning

作者:Sirui Xie, Junning Huang, Lanxin Lei, Chunxiao Liu, Zheng Ma, Wei Zhang, Liang Lin · 发表于:arXiv (Cornell University) · 年份:2018 · DOI:10.48550/arxiv.1812.09028 · 被引用次数:2 · 研究领域:Reinforcement Learning in Robotics

Reinforcement learning agents need exploratory behaviors to escape from local\noptima. These behaviors may include both immediate dithering perturbation and\ntemporally consistent exploration. To achieve these, a stochastic policy model\nthat is inherently consistent through a period of time is in desire, especially\nfor tasks with either sparse rewards or long term information. In this work, we\nintroduce a novel on-policy temporally consistent exploration strategy - Neural\nAdaptive Dropout Policy Exploration (NADPEx) - for deep reinforcement learning\nagents. Modeled as a global random variable for conditional distribution,\ndropout is incorporated to reinforcement learning policies, equipping them with\ninherent temporal consistency, even when the reward signals are sparse. Two\nfactors, gradients' alignment with the objective and KL constraint in policy\nspace, are discussed to guarantee NADPEx policy's stable improvement. Our\nexperiments demonstrate that NADPEx solves tasks with sparse reward while naive\nexploration and parameter noise fail. It yields as well or even faster\nconvergence in the standard mujoco benchmark for continuous control.\n