Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Enhancing queries for code generation with reinforcement learning

作者:Dawei Yuan, Guojun Liang, Tingting Li, Suping Liu · 发表于:Scientific Reports · 年份:2025 · DOI:10.1038/s41598-025-21271-4 · 被引用次数:2 · 研究领域:Topic Modeling、Software Engineering Research、Machine Learning and Data Classification

We present a reinforcement learning framework that enhances natural language queries to improve DeepSeek code generation. A parametric refiner (Qwen with LoRA) is trained via REINFORCE while the generator remains fixed, using a scalar reward that can combine text similarity (BLEU-4, ROUGE-L, F1, Overlap) with execution signals (unit tests, syntax/timeout penalties). On the DS1000 benchmark (800 train / 200 test), RL4QE improves the code similarity by 34.3%. Ablations show that BLEU-4 is the most reliable text reward overall (with F1 competitive on a larger scale), and LoRA with rank [Formula: see text] outperforms complete fine-tuning on most metrics while being more parameter efficient. The approach is transferred across foundation models (e.g., Qwen1.5/2/2.5 variants), where architecture often matters more than size. RL4QE is easy to integrate in practice (LoRA in attention projections) and supports reproducibility.