Scientific Writing in the Era of Large Language Models: A Computational Analysis of AI- Versus Human-Created Content
作者:Rohan Khera, Aline F Pedroso, Vipina K. Keloth, Hua Xu, Gisele Sampaio Silva, Lee H. Schwamm · 发表于:Stroke · 年份:2025 · DOI:10.1161/strokeaha.125.051913 · 被引用次数:5 · 研究领域:Artificial Intelligence in Healthcare and Education、Text Readability and Simplification、Computational and Text Analysis Methods
BACKGROUND: Large language models (LLMs) are artificial intelligence (AI) tools that can generate human expert–like content and be used to accelerate the synthesis of scientific literature, but they can spread misinformation by producing misleading content. This study sought to characterize distinguishing linguistic features in differentiating AI-generated from human-authored scientific text and evaluate the performance of AI detection tools for this task. METHODS: We conducted a computational synthesis of 34 essays on cerebrovascular topics (12 generated by large language models [Generative Pre-trained Transformer 4, Generative Pre-trained Transformer 3.5, Llama-2, and Bard] and 22 by human scientists). Each essay was rated as AI-generated or human-authored by up to 38 members of the Stroke editorial board. We compared the collective performance of experts versus GPTZero, a widely used online AI detection tool. We extracted and compared linguistic features spanning syntax (word count, complexity, and so on), semantics (polarity), readability (Flesch scores), grade level (Flesch-Kincaid), and language perplexity (or predictability) to characterize linguistic differences between AI-generated versus human-written content. RESULTS: Over 50% of the stroke experts who reviewed the study essays correctly identified 10 (83.3%) of AI-generated essays as AI, whereas they misclassified 7 (31.8%) of human-written essays as AI. GPTZero accurately classified 12 (100%) of AI-generated and ...