Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Active preference-based Gaussian process regression for reward learning and optimization

作者:Erdem Bıyık, Nicolas Huynh, Mykel J. Kochenderfer, Dorsa Sadigh · 发表于:The International Journal of Robotics Research · 年份:2023 · DOI:10.1177/02783649231208729 · 被引用次数:30 · 研究领域:Gaussian Processes and Bayesian Inference、Reinforcement Learning in Robotics、Advanced Multi-Objective Optimization Algorithms

Designing reward functions is a difficult task in AI and robotics. The complex task of directly specifying all the desirable behaviors a robot needs to optimize often proves challenging for humans. A popular solution is to learn reward functions using expert demonstrations. This approach, however, is fraught with many challenges. Some methods require heavily structured models, for example, reward functions that are linear in some predefined set of features, while others adopt less structured reward functions that may necessitate tremendous amounts of data. Moreover, it is difficult for humans to provide demonstrations on robots with high degrees of freedom, or even quantifying reward values for given trajectories. To address these challenges, we present a preference-based learning approach, where human feedback is in the form of comparisons between trajectories. We do not assume highly constrained structures on the reward function. Instead, we employ a Gaussian process to model the reward function and propose a mathematical formulation to actively fit the model using only human preferences. Our approach enables us to tackle both inflexibility and data-inefficiency problems within a preference-based learning framework. We further analyze our algorithm in comparison to several baselines on reward optimization, where the goal is to find the optimal robot trajectory in a data-efficient way instead of learning the reward function for every possible trajectory. Our results in three...