Integration of lung tissue proteomics and genome-wide association data to identify lung cancer susceptibility proteins and potential drug targets
作者:Shuai Xu, Jiajun Shi, Xie Shu, Ran Tao, Yongchao Dou, Xingyi Guo, Wanqing Wen, Yaohua Yang, B Zhang, Jie Wu, Stephen A. Deppen, Bingshan Li, Wei Zheng, Jirong Long, Qiuyin Cai · 发表于:medRxiv · 年份:2026 · DOI:10.64898/2026.06.18.26355973 · 研究领域:Genetic Associations and Epidemiology、Bioinformatics and Genomic Networks、Glutathione Transferases and Polymorphisms
Background: Proteins directly impact disease development and act as drug targets. Therefore, we integrated genomic and lung tissue proteomics data to identify lung cancer susceptibility proteins, elucidating genetic mechanisms and candidate drug targets. Method: We profiled the proteome and genome in non-neoplastic lung tissue from 200 lung cancer patients. Using this data, we constructed genetic models to predict abundance across the proteome in lung tissue. We applied these models to genome-wide association study (GWAS) data from 55,174 lung cancer cases and 1,294,174 controls to evaluate their associations with the risk of lung cancer, overall and by major histological subtypes. Bayesian colocalization and Mendelian randomization (MR) analyses were used to prioritize putative causal proteins, which were cross-referenced with three main drug-protein databases to identify potential therapeutic targets. Results: We identified 29 proteins associated with lung cancer risk at a false discovery rate < 5%, including 25 for overall lung cancer, two (AQP3 and IL18) specifically for adenocarcinoma, and another two (HMGN2 and HLA-DMB) for squamous cell carcinoma. Of them, genes encoding 17 proteins reside at least 2Mb away from any known GWAS risk loci, including 14 for overall lung cancer (HYI, GPX1, GMPPB, DSP, HDDC2, MTCH2, SUOX, JMJD7, PDIA3, IL16, IQGAP1, SULT1A2, ARHGAP27, and TYMP) and three for subtypes (AQP3, IL18, and HMGN2). Among the 12 proteins located within the known ri...