Scholay

学术搜索 · AI 审稿 · LaTeX 协作

HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

作者:Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie, Ashmal Vayani, Ahmed Y. Radwan, M. S. Chettiar, Amandeep Singh, Mubarak Shah, D. Pandya · 发表于:ACM Transactions on Intelligent Systems and Technology · 年份:2025 · DOI:10.1145/3845999 · 被引用次数:24 · 研究领域:Computer Science

Although recent large multimodal models (LMMs) show impressive progress on vision–language tasks, their alignment with human-centered (HC) principles such as fairness, ethics, inclusivity, empathy, and robustness is often overlooked. Existing LMM benchmarks are largely accuracy-agnostic. We present HumaniBench, a unified framework for characterizing HC alignment across realistic, socially grounded visual contexts. It contains 32,000 expert-verified image–question pairs from real-world news imagery, each mapped to one or more HC principles through explicit metrics. Comparing 15 state-of-the-art LMMs reveals consistent trade-offs: proprietary systems lead on ethics, reasoning, and empathy, while open-source models show superior visual grounding and resilience. All models show persistent gaps in fairness and multilingual inclusivity. Chain-of-thought prompting and test-time scaling yield 8–12% gains on several HC dimensions. HumaniBench enables fine-grained analysis of alignment trade-offs not captured by conventional multimodal benchmarks. Project: https://vectorinstitute.github.io/humanibench/ Data: https://huggingface.co/vector-institute/HumaniBench Code: https://github.com/VectorInstitute/HumaniBench