Scholay

学术搜索 · AI 审稿 · LaTeX 协作

PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

作者:Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, Cheng Yu · 年份:2025 · DOI:10.18653/v1/2025.acl-long.1230 · 被引用次数:7 · 研究领域:Business Process Modeling and Analysis

Process-level Reward Models (PRMs) are crucial for complex reasoning and decisionmaking tasks, where each intermediate step plays an important role in the reasoning process.Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios.However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically.To address this gap, we introduce PRMBENCH, a processlevel benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs.PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity.In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs.These findings reveal the challenges inherent in processlevel evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development.