Data from Design and Creation of a Racially Diverse Lung Cancer Registry with Detailed Genomic and Environmental Annotation
作者:Luchang Cui, Juhong Lee, J S Miller, Minmeng Tang, Sajjad Abedian, Nasser K. Altorki, Robert Crupi, Lauren K. Groner, Neal I. Lindeman, Laura C. Pinheiro, Rulla M. Tamimi, Jonathan Villena‐Vargas, Anil Vachani, Julie A. Barta, H. Oliver Gao, Evan Sholle, James P. Solomon, Christine Garcia, Eunji Choi, Yiwey Shieh · 年份:2026 · DOI:10.1158/1055-9965.c.8547435 · 研究领域:Lung Cancer Treatments and Mutations、Lung Cancer Diagnosis and Treatment、Ferroptosis and cancer prognosis
<div>AbstractBackground:<p>The proportion of lung cancers affecting individuals who have never smoked is growing, with these cancers being prone to harbor mutations in the <i>EGFR</i> gene. Little is known about risk factors and prognostic indicators for <i>EGFR</i>-mutant cancers, with current research limited by the scarcity of datasets integrating genomic, clinical, and environmental data.</p>Methods:<p>We created the Meyer Cancer Center Molecularly Enhanced Lung Cancer Database (MCC-MELD), including lung cancer cases from a large catchment area in New York City. We identified cases through linkage to our institution’s cancer registry and a clinician-initiated, manually curated database. We linked all cases to the electronic health record and in-house tumor genomic testing results. We used natural language processing (NLP) to extract unstructured genomic testing results and detailed smoking history. We linked geocoded addresses to detailed area-level measures.</p>Results:<p>MCC-MELD contains 9,573 patients with lung cancer diagnosed from 1988 to 2024, of whom 20% were non-Hispanic Asian, 14% were non-Hispanic Black, and 8% were Hispanic. We identified 1,092 (11.4%) <i>EGFR</i>-mutant cancers, with NLP identifying 397 cases not identified by structured data. NLP showed high accuracy in ascertaining <i>EGFR</i> status (97%) and quantitative smoking history variables (90%–98%). Never smokers m...