Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Development of a Natural Language Processing Model for Extracting Kidney Biopsy Pathology Diagnoses

作者:Shane A. Bobart, Enshuo Hsu, Thomas Potter, Luan D. Truong, Amy D. Waterman, Stephen Jones, Tariq Shafi · 发表于:Kidney Medicine · 年份:2025 · DOI:10.1016/j.xkme.2025.101047 · 被引用次数:4 · 研究领域:AI in cancer detection、Artificial Intelligence in Healthcare and Education、Machine Learning in Healthcare

Rationale & Objective: Kidney biopsy reports are in a nonindexed text format, and the diagnosis requires labor-intensive manual abstraction. Natural language processing (NLP) has not been rigorously tested for kidney biopsy diagnosis extraction. Our objective was to develop an accurate model to extract the biopsy diagnosis from free-text reports. Study Design: Text classification using NLP. Setting & Participants: 2,666 patients with 3,042 native kidney biopsy reports in the Portable Document Format, from June 2016 to December 2023. Predictor: Kidney biopsy diagnosis. Outcomes: The performance of the NLP algorithm for all and the 20 most common diagnoses based on precision, recall, F1 score, and area under the receiver operating curve (AUROC). Analytical Approach: A domain expert manually abstracted the diagnosis, and a renal pathologist validated a random subset (n = 200). Structured Query Language server and Python processed reports into machine-readable free text. We used PubMed Bidirectional Encoder Representations from Transformers to develop our NLP algorithm. We randomly split the reports into training (80%; n = 2,434) and testing (20%; n = 608) sets to train the NLP system. We further divided the testing set into 20% validation and 80% fine-tuning sets. Results: The median age was 57 years, with 50% female, 29% African Americans, and 23% Hispanic participants. The 5 most frequent glomerular diagnoses were diabetic kidney disease (23.7%), focal segmental glomeruloscler...