Scholay

学术搜索 · AI 审稿 · LaTeX 协作

NextPolish: a fast and efficient genome polishing tool for long-read assembly

作者:Jiang Hu, Junpeng Fan, Zongyi Sun, Shanlin Liu · 发表于:Bioinformatics · 年份:2019 · DOI:10.1093/bioinformatics/btz891 · 被引用次数:1333 · 研究领域:Genomics and Phylogenetic Studies、Genome Rearrangement Algorithms、CRISPR and Genetic Engineering

MOTIVATION: Although long-read sequencing technologies can produce genomes with long contiguity, they suffer from high error rates. Thus, we developed NextPolish, a tool that efficiently corrects sequence errors in genomes assembled with long reads. This new tool consists of two interlinked modules that are designed to score and count K-mers from high quality short reads, and to polish genome assemblies containing large numbers of base errors. RESULTS: When evaluated for the speed and efficiency using human and a plant (Arabidopsis thaliana) genomes, NextPolish outperformed Pilon by correcting sequence errors faster, and with a higher correction accuracy. AVAILABILITY AND IMPLEMENTATION: NextPolish is implemented in C and Python. The source code is available from https://github.com/Nextomics/NextPolish. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.