AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
作者:Shaona Ghosh, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, Sean McGregor, Kenneth Fricklas, Mala Kumar, Quentin Feuillade--Montixi, Kurt Bollacker, Felix Friedrich, Ryan Tsang, Bertie Vidgen, Alicia Parrish, Chris Knotz, Eleonora Presani, Jonathan Bennion, Marisa Ferrara Boston, Mike Kuniavsky, W. Hutiri, James Ezick, Malek Ben Salem, Rajat Sahay, Sujata Goswami, Usman Gohar, Ben Huang, Supheakmungkol Sarin, Elie Alhajjar, Canyu Chen, Roman Eng, K. Manjusha, Virendra Mehta, Eileen Long, M. Emani, Natan Vidra, Benjamin Rukundo, Abolfazl Shahbazi, Kongtao Chen, Rajat Ghosh, Vithursan Thangarasa, Pierre Peigné, Abhinavkumar Singh, Max Bartolo, Satyapriya Krishna, Mubashara Akhtar, R. Gold, C. Coleman, Luis Oala, Vassil Tashev, Joseph Marvin Imperial, Amy Russ, Sasidhar Kunapuli, Nicolas Miailhe, Julien Delaunay, Bhaktipriya Radharapu, Rajat C. Shinde, Tuesday, Debojyoti Dutta, D. Grabb, Ananya Gangavarapu, Saurav Sahay, Agasthya Gangavarapu, P. Schramowski, S. Singam, Tom David, Xudong Han, P. Mammen, Tarunima Prabhakar, Venelin Kovatchev, Ahmed M. Ahmed, Kelvin N. Manyeki, Sandeep Madireddy, F. Khomh, Fedor Zhdanov, Joachim Baumann, N. Vasan, Xianjun Yang, Carlos Mougn, J. Varghese, H. Chinoy, Seshakrishna Jitendar, M. Maskey, C. Hardgrove, Tianhao Li, Aakash Gupta, Emil Joswin, Yifan Mai, Shachi H. Kumar, Çigdem Patlak, K. Lu, Vincent Alessi, Sree Bhargavi Balija, Chenhe Gu, R. Sullivan, J. Gealy, Matt Lavrisa, James Goel, Peter Mattson, Percy Liang, Joaquin Vanschoren · 发表于:arXiv.org · 年份:2025 · DOI:10.48550/arXiv.2503.05731 · 被引用次数:35 · 研究领域:Computer Science
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehensive industry-standard benchmark for assessing AI-product risk and reliability. Its development employed an open process that included participants from multiple fields. The benchmark evaluates an AI system's resistance to prompts designed to elicit dangerous, illegal, or undesirable behavior in 12 hazard categories, including violent crimes, nonviolent crimes, sex-related crimes, child sexual exploitation, indiscriminate weapons, suicide and self-harm, intellectual property, privacy, defamation, hate, sexual content, and specialized advice (election, financial, health, legal). Our method incorporates a complete assessment standard, extensive prompt datasets, a novel evaluation framework, a grading and reporting system, and the technical as well as organizational infrastructure for long-term support and evolution. In particular, the benchmark employs an understandable five-tier grading scale (Poor to Excellent) and incorporates an innovative entropy-based system-response evaluation. In addition to unveiling the benchmark, this report also identifies limitations of our method and of building safety benchmarks generally, including evaluator uncertainty and the constraints of single-turn interactions. This work represents a crucial step toward establishing global standards for AI risk and reliability e...