Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Deterministic delay-aware reinforcement learning

作者:Sathira Dilshan Bataduwaarachchi, Zoran Najdovski, Hieu Trinh, Chee Peng Lim, Van Thanh Huynh · 发表于:Robotics and Autonomous Systems · 年份:2025 · DOI:10.1016/j.robot.2025.105271 · 被引用次数:1 · 研究领域:Reinforcement Learning in Robotics、Age of Information Optimization、Neural Networks and Reservoir Computing

Reinforcement Learning (RL) has effectively paved the way in achieving robotic control during the past decade. As a result, the avenue of integrating RL-powered robotic control and teleoperation has caught the attention of researchers. Every RL framework involves the basis of suitable observation and action communication between the environment and the agent, and the involvement of teleoperation can introduce random time delays within the said communication process. Achieving robotic control under such constraints remains an untapped area in the domain of reinforcement learning. We take the initiative to achieve the goal of robotic control while handling delays in the RL setting based on a fitting Markov Decision Process (MDP) structure. Our algorithm will learn a deterministic policy and can tackle control environments, especially robotic manipulation environments, using observations with proprioceptive information. We methodically present the theoretical adjustments based on an existing dominant off-policy algorithm to express the algorithm’s competency with proof of convergence. We perform experimentations with DeepMind Control Suite, illustrating significant results showing the algorithm’s capabilities in learning complex environments powered by delay-aware RL.