Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Towards Automatic Boundary Detection for Human-AI Collaborative Hybrid Essay in Education

作者:Zijie Zeng, Lele Sha, Yuheng Li, Kaixun Yang, Dragan Gašević, Guangliang Chen · 发表于:Proceedings of the AAAI Conference on Artificial Intelligence · 年份:2024 · DOI:10.1609/aaai.v38i20.30258 · 被引用次数:16 · 研究领域:Online Learning and Analytics、Advanced Data Processing Techniques

The recent large language models (LLMs), e.g., ChatGPT, have been able to generate human-like and fluent responses when provided with specific instructions. While admitting the convenience brought by technological advancement, educators also have concerns that students might leverage LLMs to complete their writing assignments and pass them off as their original work. Although many AI content detection studies have been conducted as a result of such concerns, most of these prior studies modeled AI content detection as a classification problem, assuming that a text is either entirely human-written or entirely AI-generated. In this study, we investigated AI content detection in a rarely explored yet realistic setting where the text to be detected is collaboratively written by human and generative LLMs (termed as hybrid text for simplicity). We first formalized the detection task as identifying the transition points between human-written content and AI-generated content from a given hybrid text (boundary detection). We constructed a hybrid essay dataset by partially and randomly removing sentences from the original student-written essays and then instructing ChatGPT to fill in for the incomplete essays. Then we proposed a two-step detection approach where we (1) separated AI-generated content from human-written content during the encoder training process; and (2) calculated the distances between every two adjacent prototypes (a prototype is the mean of a set of consecutive senten...