Stigmatizing Language in Large Language Models for Alcohol and Substance Use Disorders: A Multimodel Evaluation and Prompt Engineering Approach.
作者:Yichen Wang, K. Hsu, C. Brokus, Yuting Huang, Nneka N. Ufere, Sarah E Wakeman, James Zou, Wei Zhang · 发表于:Journal of addiction medicine · 年份:2025 · DOI:10.1097/adm.0000000000001536 · 被引用次数:7 · 研究领域:Medicine
OBJECTIVES Large language models (LLMs) are increasingly used in health care communication but can inadvertently perpetuate stigmatizing language toward individuals with alcohol and substance use disorders. Despite growing interest in LLM performance, a focused evaluation of their propensity for SL and strategies to mitigate it remains lacking. METHODS We generated 60 clinically relevant questions ["prompts"; 20 each for alcohol use disorder (AUD), alcohol-associated liver disease (ALD), and substance use disorder (SUD)] and tested 14 LLMs. Two physicians independently assessed all responses for stigmatizing language using guidelines from the National Institute on Drug Abuse and the National Institute on Alcohol Abuse and Alcoholism; discrepancies were resolved by a third physician. We employed iterative prompt engineering (PE)-a process of strategically crafting input instructions to guide model outputs towards nonstigmatizing language-to reduce stigmatizing language by incorporating a list of specific terms to avoid and identifying model-specific pitfalls. We compared the prevalence of SL in responses to native prompts (baseline, unengineered) versus engineered prompts, adjusting for word count in multivariate analyses. RESULTS Of 840 responses generated from native prompts, 297 (35.4%) contained stigmatizing language, totaling 592 terms. With prompt engineering, only 53 (6.3%) of 840 responses contained stigmatizing language, comprising 104 terms. Prompts on topic of A...