The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model
作者:Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, Matthieu Geist, Yuejie Chi · 发表于:Operations Research · 年份:2026 · DOI:10.1287/opre.2025.2240 · 被引用次数:3 · 研究领域:Reinforcement Learning in Robotics、Adversarial Robustness in Machine Learning
Robust Reinforcement Learning Without Compromising Data Efficiency This paper investigates model robustness in reinforcement learning (RL) to reduce the sim-to-real gap in practice. We adopt the framework of distributionally robust Markov decision processes (RMDPs), which aims to learn a policy that optimizes worst-case performance over a prescribed uncertainty set around a nominal Markov decision process (MDP). Despite recent efforts, the sample complexity of RMDPs has remained largely unresolved, leaving open whether distributional robustness has any statistical consequences when benchmarked against standard RL. Assuming access to a generative model of the nominal MDP, the paper provides a near-optimal characterization of the sample complexity of RMDPs across the full range of uncertainty levels under two common uncertainty sets specified by either total variation (TV) distance or chi-squared divergence. Somewhat surprisingly, the results reveal that RMDPs are not necessarily easier or harder to learn than standard MDPs. The statistical consequences of the robustness requirement depend heavily on the size and shape of the uncertainty set, requiring less data in the TV case but more in the chi-squared case compared with standard MDPs.