Scholay

学术搜索 · AI 审稿 · LaTeX 协作

ACECODER: Acing Coder RL via Automated Test-Case Synthesis

作者:Huaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie, Xiaotong Chen, Wenhu Chen · 年份:2025 · DOI:10.18653/v1/2025.acl-long.587 · 被引用次数:4 · 研究领域:Software Testing and Debugging Techniques、Model-Driven Software Engineering Techniques、Real-time simulation and control systems

Most progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexplored, primarily due to the lack of reliable reward data/model in the code domain.In this paper, we address this challenge by leveraging automated large-scale testcase synthesis to enhance code model training.Specifically, we design a pipeline that generates extensive (question, test-cases) pairs from existing code data.Using these test cases, we construct preference pairs based on pass rates over sampled programs to train reward models with Bradley-Terry loss.It shows an average of 10-point improvement for Llama-3.1-8B-Ins and 5-point improvement for Qwen2.5-Coder-7B-Insthrough best-of-32 sampling, making the 7B model on par with 236B DeepSeek-V2.5.Furthermore, we conduct reinforcement learning with both reward models and testcase pass rewards, leading to consistent improvements across HumanEval, MBPP, Big-CodeBench, and LiveCodeBench (V4).Notably, we follow the R1-style training to start from Qwen2.5-Coder-base directly and show that our RL training can improve model on HumanEval-plus by over 25% and MBPP-plus by 6% for merely 80 optimization steps.We believe our results highlight the huge potential of reinforcement learning in coder models.