Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Automated Extraction of Mortality Information From Publicly Available Sources Using Large Language Models: Development and Evaluation Study

作者:Mohammed Ali Al-Garadi, Michele L. Lenoue-Newton, Michael E. Matheny, Melissa L McPheeters, Jill M Whitaker, Jeff Deere, Michael F McLemore, Dax Westerman, Mirza S. Khan, José J. Hernández‐Muñoz, Xi Wang, Aida Kuzucan, Rishi Desai, Ruth Reeves · 发表于:Journal of Medical Internet Research · 年份:2025 · DOI:10.2196/71113 · 被引用次数:5 · 研究领域:Data-Driven Disease Surveillance、Data Quality and Management、Mental Health via Writing

Background: Mortality is a critical variable in health care research, especially for evaluating medical product safety and effectiveness. However, inconsistencies in the availability and timeliness of death date and cause of death (CoD) information present significant challenges. Conventional sources such as the National Death Index and electronic health records often experience data lags, missing fields, or incomplete coverage, limiting their utility in time-sensitive or large-scale studies. With the growing use of social media, crowdfunding platforms, and web-based memorials, publicly available digital content has emerged as a potential supplementary source for mortality surveillance. Despite this potential, accurate tools for extracting mortality information from such unstructured data sources remain underdeveloped. Objective: The aim of the study is to develop scalable approaches using natural language processing (NLP) and large language models (LLMs) for the extraction of mortality information from publicly available web-based data sources, including social media platforms, crowdfunding websites, and web-based obituaries, and to evaluate their performance across various sources. Methods: Data were collected from public posts on X (formerly known as Twitter), GoFundMe campaigns, memorial websites (EverLoved and TributeArchive), and web-based obituaries from 2015 to 2022, focusing on US-based content relevant to mortality. We developed an NLP pipeline using transformer-bas...