Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Reinforcement Learning with Factored States and Actions

作者:Brian Sallans, Geoffrey E. Hinton · 年份:2004 · DOI:10.5555/1005332.1016794 · 被引用次数:138 · 研究领域:Reinforcement Learning in Robotics、Data Stream Mining Techniques、Bayesian Modeling and Causal Inference

A novel approximation method is presented for approximating the value function and selecting good actions for Markov decision processes with large state and action spaces. The method approximates state-action values as negative free energies in an undirected graphical model called a product of experts. The model parameters can be learned efficiently because values and derivatives can be efficiently computed for a product of experts. Actions can be found even in large factored action spaces by the use of Markov chain Monte Carlo sampling. Simulation results show that the product of experts approximation can be used to solve large problems. In one simulation it is used to find actions in action spaces of size 2 40.