SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
作者:Yazhou Zhang, Chunwang Zou, Zheng Lian, Prayag Tiwari, Jing Qin · 发表于:IEEE Transactions on Affective Computing · 年份:2025 · DOI:10.1109/taffc.2025.3604806 · 被引用次数:15 · 研究领域:Wildlife Ecology and Conservation
In the era of large language models (LLMs), tasks associated with “System I” cognition—those that are fast, automatic, and intuitive, such as sentiment analysis and text classification—are often considered effectively solved. However, sarcasm remains a persistent challenge. As a subtle and complex linguistic phenomenon, sarcasm frequently involves rhetorical devices such as hyperbole and figurative language to express implicit sentiments and intentions, demanding a higher level of abstraction and pragmatic reasoning than standard sentiment analysis. This raises concerns about whether current claims of LLM success extend robustly to the domain of sarcasm understanding. To systematically investigate this issue, we introduce a new high-quality multi-modal sarcasm detection dataset, termedAMSD, and construct a comprehensive evaluation benchmark,SarcasmBench. Our benchmark encompasses 16 state-of-the-art (SOTA) LLMs and 8 strong pretrained language models (PLMs), evaluated across six widely-used textual sarcasm datasets and three multi-modal sarcasm benchmarks. We adopt three popular prompting paradigms: zero-shot input/output (IO) prompting, few-shot IO prompting, and chain-of-thought (CoT) prompting. Our extensive experiments yield three key findings: (1) current LLMs underperform supervised PLMs based sarcasm detection baselines. This suggests that significant efforts are still required to improve LLMs' understanding of human sarcasm. (2) GPT-4 and Gemini 2.0 consistently and s...