Multi-Criteria Evaluation of Large Language Models (LLMs): Balancing Performance and Security
作者:Daniel Mendonça Colares, Plácido Rogério Pinheiro, R. A. Freitas Filho · 发表于:IEEE Access · 年份:2026 · DOI:10.1109/access.2026.3665546 · 被引用次数:1 · 研究领域:Artificial Intelligence in Healthcare and Education、Web Application Security Vulnerabilities、Computational and Text Analysis Methods
Because of its functionality and practicality, Large Language Models (LLMs) have been widely discussed, with a large number of benchmarking being done to evaluate them, especially their efficiency. But despite their numerous applications and the significant benefits they offer, LLMs have proven to be extremely susceptible to attacks of various natures due to their, often unknown, large number of vulnerabilities, characteristics often ignored by benchmarking. Given that, this paper aims to develop a multi-criteria method to assist stakeholders in selecting the most suitable Large Language Model taking into account based on both its efficiency in carrying out tasks of various natures, such as math and reasoning, and its capability to resist a large range of security vulnerabilities, such as prompt injection and jailbreaking. This study utilized the Analytic Hierarchy Process (AHP) along with tools developed to evaluate the capabilities of LLMs in multi-interaction dialogues and LLM vulnerability scanner applied in open source models. The analysis showed that an more efficient model does not mean that it is safer. In addition, it reveals an efficient method for analyzing both model performance and security issues.