Phenonizer: A Fine‐Grained Phenotypic Named Entity Recognizer for Chinese Clinical Texts
作者:Qunsheng Zou, Kuo Yang, Zixin Shu, Kai Chang, Qiguang Zheng, Yi Zheng, Kezhi Lu, Ning Xu, Haoyu Tian, Xiaomeng Li, Yuxia Yang, Yana Zhou, Haibin Yu, Xiao–Ping Zhang, Jianan Xia, Qiang Zhu, Josiah Poon, Simon Poon, Runshun Zhang, Xiaodong Li, Xuezhong Zhou · 发表于:BioMed Research International · 年份:2022 · DOI:10.1155/2022/3524090 · 被引用次数:10 · 研究领域:Topic Modeling、Natural Language Processing Techniques、Biomedical Text Mining and Ontologies
Biomedical named entity recognition (BioNER) from clinical texts is a fundamental task for clinical data analysis due to the availability of large volume of electronic medical record data, which are mostly in free text format, in real‐world clinical settings. Clinical text data incorporates significant phenotypic medical entities (e.g., symptoms, diseases, and laboratory indexes), which could be used for profiling the clinical characteristics of patients in specific disease conditions (e.g., Coronavirus Disease 2019 (COVID‐19)). However, general BioNER approaches mostly rely on coarse‐grained annotations of phenotypic entities in benchmark text dataset. Owing to the numerous negation expressions of phenotypic entities (e.g., “no fever,” “no cough,” and “no hypertension”) in clinical texts, this could not feed the subsequent data analysis process with well‐prepared structured clinical data. In this paper, we developed Human‐machine Cooperative Phenotypic Spectrum Annotation System ( http://www.tcmai.org/login , HCPSAS) and constructed a fine‐grained Chinese clinical corpus. Thereafter, we proposed a phenotypic named entity recognizer: Phenonizer, which utilized BERT to capture character‐level global contextual representation, extracted local contextual features combined with bidirectional long short‐term memory, and finally obtained the optimal label sequences through conditional random field. The results on COVID‐19 dataset show that Phenonizer outperforms those methods based...