Model-Agnostic Empirical Evaluation of Test-Driven Prompt Engineering on Improving Accuracy and Efficiency in Large Language Models Python Code Generation
作者:Muhammad Rizqullah, Emad Albassam · 发表于:IEEE Access · 年份:2026 · DOI:10.1109/access.2026.3662817 · 被引用次数:1 · 研究领域:Computer Science
Although Large Language Models (LLMs) are widely adopted for Python code generation, the generated code can be semantically incorrect, requiring iterations of evaluation and refinement. Test-driven prompt (TDP) engineering systematically incorporates test cases within the prompt structure for code generation task. However, existing research lacks a comprehensive evaluation of TDP engineering across diverse LLMs, programming problems, and associated explainability analysis. In this paper, we propose a specification clarification mechanism to explain TDP’s effectiveness in generating Python code while also uncovering novel insights through multimodel analysis. To this end, we created an experimental framework that facilitates evaluation of TDP engineering against normal prompting across first-attempt and remediation scenarios using eight models (GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet/Haiku, Qwen 32B/14B/7B/3B Coder) and three benchmarks (HumanEval, MBPP, Code Contests). Qualitative analysis of code produced with TDP for programming problems of different difficulty shows consistent accuracy gains, from following formatting specifications to offering safeguards against logical errors, where each level of difficulty shows its own main advantage. Quantitatively, TDP achieves superior aggregate performance in all of 16 model-dataset combinations, with 6.09% average improvement (95% CI: [4.01, 8.18], p ¡ 0.0001, Cohen’s d = 1.08). A key practical finding is that TDP enables democrati...