Scholay

学术搜索 · AI 审稿 · LaTeX 协作

LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems

作者:Tianyang Duan, Zongyuan Zhang, Zheng Lin, Songxiao Guo, Xiuxian Guan, Guangyu Wu, Zihan Fang, Haotian Meng, Xia Du, Jizhe Zhou, Heming Cui, Jun Luo, Yue Gao · 发表于:IEEE Transactions on Mobile Computing · 年份:2026 · DOI:10.1109/tmc.2026.3682994 · 研究领域:Reinforcement Learning in Robotics、Advanced Bandit Algorithms Research、Distributed Control Multi-Agent Systems

Multi-agent reinforcement learning (MARL) has been increasingly adopted in many real-world applications. While MARL enables decentralized deployment on resource-constrained edge devices, it suffers from severe non-stationarity due to the synchronous updates of agent policies. This non-stationarity results in unstable training and poor policy convergence, especially as the number of agents increases. In this paper, we propose Refinement-Enhanced LLM Expert Demonstrations (RELED), a scalable MARL framework that integrates large language model (LLM)-driven expert demonstrations with autonomous agent exploration. RELED incorporates a Stationarity-Aware Expert Demonstration module, which leverages theoretical non-stationarity bounds to enhance the quality of LLM-generated expert trajectories, thus providing high-reward and training-stable samples for each agent. Moreover, a Hybrid Expert-Agent Policy Optimization module adaptively balances each agent's learning from both expert-generated and agent-generated trajectories, accelerating policy convergence and improving generalization. Extensive experiments with real city networks based on OpenStreetMap demonstrate that RELED achieves superior performance compared to state-of-the-art MARL methods.