Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Benchmarking LLMs Against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

作者:Jianhao Yan, Pingchuan Yan, Yulong Chen, Jing Li, Xianchao Zhu, Yue Zhang · 发表于:IEEE Transactions on Big Data · 年份:2025 · DOI:10.1109/tbdata.2025.3644594 · 被引用次数:6 · 研究领域:Natural Language Processing Techniques、Topic Modeling、Text Readability and Simplification

This study presents a comprehensive evaluation of the translation capabilities of existing LLMs, such as GPT-4, ALMA-R, and Deepseek-R1, compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three language pairs (Chinese$\longleftrightarrow$English, Russian$\longleftrightarrow$English, and Chinese$\longleftrightarrow$Hindi) and three domains (News, Technology, and Biomedical). Our findings reveal that LLMs achieve performance comparable to junior-level translators in terms of total errors, while still lagging behind senior translators. Unlike traditional Neural Machine Translation systems, which show significant performance degradation in resource-poor language directions, LLMs like GPT-4 maintain consistent translation quality across all evaluated language pairs. Through qualitative analysis, we identify distinctive patterns in translation approaches: GPT-4 tends toward overly literal translations and exhibits lexical inconsistency, while human translators sometimes over-interpret context and introduce hallucinations. This study presents a systematic comparison between LLMs and human translators across different proficiency levels, providing valuable insights into the current capabilities and limitations of LLM-based translation systems.