Scholay

学术搜索 · AI 审稿 · LaTeX 协作

VirFinder: a novel k-mer based tool for identifying viral sequences from assembled metagenomic data

作者:Jie Ren, Nathan A. Ahlgren, Yang Young Lu, Jed A. Fuhrman, Fengzhu Sun · 发表于:Microbiome · 年份:2017 · DOI:10.1186/s40168-017-0283-5 · 被引用次数:732 · 研究领域:Bacteriophages and microbial interactions、Animal Virus Infections Studies、Genomics and Phylogenetic Studies

BACKGROUND: Identifying viral sequences in mixed metagenomes containing both viral and host contigs is a critical first step in analyzing the viral component of samples. Current tools for distinguishing prokaryotic virus and host contigs primarily use gene-based similarity approaches. Such approaches can significantly limit results especially for short contigs that have few predicted proteins or lack proteins with similarity to previously known viruses. METHODS: We have developed VirFinder, the first k-mer frequency based, machine learning method for virus contig identification that entirely avoids gene-based similarity searches. VirFinder instead identifies viral sequences based on our empirical observation that viruses and hosts have discernibly different k-mer signatures. VirFinder's performance in correctly identifying viral sequences was tested by training its machine learning model on sequences from host and viral genomes sequenced before 1 January 2014 and evaluating on sequences obtained after 1 January 2014. RESULTS: VirFinder had significantly better rates of identifying true viral contigs (true positive rates (TPRs)) than VirSorter, the current state-of-the-art gene-based virus classification tool, when evaluated with either contigs subsampled from complete genomes or assembled from a simulated human gut metagenome. For example, for contigs subsampled from complete genomes, VirFinder had 78-, 2.4-, and 1.8-fold higher TPRs than VirSorter for 1, 3, and 5 kb contigs,...