Sorting Through the Safety Data Haystack: Using Machine Learning to Identify Individual Case Safety Reports in Social-Digital Media
作者:Shaun Comfort, Sujan Perera, Zoë Hudson, Darren Dorrell, Shawman Meireis, Meenakshi Nagarajan, C. R. Ramakrishnan, Jennifer Fine · 发表于:Drug Safety · 年份:2018 · DOI:10.1007/s40264-018-0641-7 · 被引用次数:50 · 研究领域:Pharmacovigilance and Adverse Drug Reactions、Biomedical Text Mining and Ontologies、Drug-Induced Adverse Reactions
INTRODUCTION: There is increasing interest in social digital media (SDM) as a data source for pharmacovigilance activities; however, SDM is considered a low information content data source for safety data. Given that pharmacovigilance itself operates in a high-noise, lower-validity environment without objective 'gold standards' beyond process definitions, the introduction of large volumes of SDM into the pharmacovigilance workflow has the potential to exacerbate issues with limited manual resources to perform adverse event identification and processing. Recent advances in medical informatics have resulted in methods for developing programs which can assist human experts in the detection of valid individual case safety reports (ICSRs) within SDM. OBJECTIVE: In this study, we developed rule-based and machine learning (ML) models for classifying ICSRs from SDM and compared their performance with that of human pharmacovigilance experts. METHODS: We used a random sampling from a collection of 311,189 SDM posts that mentioned Roche products and brands in combination with common medical and scientific terms sourced from Twitter, Tumblr, Facebook, and a spectrum of news media blogs to develop and evaluate three iterations of an automated ICSR classifier. The ICSR classifier models consisted of sub-components to annotate the relevant ICSR elements and a component to make the final decision on the validity of the ICSR. Agreement with human pharmacovigilance experts was chosen as the pr...