Extracting Concepts for Precision Oncology from the Biomedical\n Literature
作者:Nicholas Greenspan, Yuqi Si, Kirk Roberts · 发表于:arXiv (Cornell University) · 年份:2020 · DOI:10.48550/arxiv.2010.00074 · 被引用次数:1 · 研究领域:Biomedical Text Mining and Ontologies、Topic Modeling、Natural Language Processing Techniques
This paper describes an initial dataset and automatic natural language\nprocessing (NLP) method for extracting concepts related to precision oncology\nfrom biomedical research articles. We extract five concept types: Cancer,\nMutation, Population, Treatment, Outcome. A corpus of 250 biomedical abstracts\nwere annotated with these concepts following standard double-annotation\nprocedures. We then experiment with BERT-based models for concept extraction.\nThe best-performing model achieved a precision of 63.8%, a recall of 71.9%, and\nan F1 of 67.1. Finally, we propose additional directions for research for\nimproving extraction performance and utilizing the NLP system in downstream\nprecision oncology applications.\n