Repeat and haplotype aware error correction in nanopore sequencing reads with DeChat
作者:Yuansheng Liu, Yichen Li, Enlian Chen, Jialu Xu, Wenhai Zhang, Xiangxiang Zeng, Xiao Luo · 发表于:Communications Biology · 年份:2024 · DOI:10.1038/s42003-024-07376-y · 被引用次数:21 · 研究领域:Genomics and Phylogenetic Studies、Molecular Biology Techniques and Applications、Environmental DNA in Biodiversity Studies
Error self-correction is crucial for analyzing long-read sequencing data, but existing methods often struggle with noisy data or are tailored to technologies like PacBio HiFi. There is a gap in methods optimized for Nanopore R10 simplex reads, which typically have error rates below 2%. We introduce DeChat, a novel approach designed specifically for these reads. DeChat enables repeat- and haplotype-aware error correction, leveraging the strengths of both de Bruijn graphs and variant-aware multiple sequence alignment to create a synergistic approach. This approach avoids read overcorrection, ensuring that variants in repeats and haplotypes are preserved while sequencing errors are accurately corrected. Benchmarking on simulated and real datasets shows that DeChat-corrected reads have significantly fewer errors—up to two orders of magnitude lower—compared to other methods, without losing read information. Furthermore, DeChat-corrected reads clearly improves genome assembly and taxonomic classification. DeChat is a novel error correction tool for Nanopore R10 reads (<2% error), combining de Bruijn graphs and variant-aware alignment to preserve repeats and haplotypes while reducing errors up to 100x, outperforming existing methods.