Boosting the Power of Rare Variant Association Studies by Imputation Using Large-scale Sequencing Population
作者:Jinglan Dai, Yixin Zhang, Yuan Gao, Hongru Li, Sha Du, Hao Hong, Dongfang You, Zaiming Li, Ruyang Zhang, Yang Zhao, Zhonghua Liu, David C. Christiani, Feng Chen, Sipeng Shen · 发表于:Genomics Proteomics & Bioinformatics · 年份:2025 · DOI:10.1093/gpbjnl/qzaf084 · 被引用次数:3 · 研究领域:Genetic Associations and Epidemiology、Epigenetics and DNA Methylation、Genetic Syndromes and Imprinting
With the emergence of population-scale whole-genome sequencing (WGS), rare variants can be captured precisely. Studying rare variants explains part of the heritability of complex traits that is overlooked by conventional genome-wide association studies (GWASs). However, the extent to which imputed data can approximate or improve upon the power of WGS data in rare variant association studies remains unclear. Using the UK Biobank WGS data (n = 150,119) as the ground truth, we first evaluated the consistency of rare variants in the single-nucleotide polymorphism (SNP) array data imputed using TOPMed or HRC+UK10K reference panel. Imputation quality (average R2) of the TOPMed-imputed data reached 0.6 even for extremely rare variants with minor allele count ≤ 5. TOPMed-imputed data were closer to WGS data across three ethnic groups, with average Cramer's V > 0.75. Furthermore, association tests were performed on 45 traits. At the same sample size (n = 150,119), neither imputed dataset outperformed WGS data, but the results of the TOPMed-imputed data were more consistent with those of WGS data. When the sample size was increased to 488,377, the number of significant rare variants identified from the TOPMed-imputed data increased by 27.71% for quantitative traits and by approximately 10-fold for binary traits. Finally, we meta-analyzed the association results of SNP array and WGS for lung cancer and epithelial ovarian cancer, respectively. Compared to WGS-based results, more signific...