Scholay

学术搜索 · AI 审稿 · LaTeX 协作

VADR: validation and annotation of virus sequence submissions to GenBank

作者:Alejandro A. Schäffer, Eneida Hatcher, Linda Yankie, Lara Shonkwiler, J. Rodney Brister, Ilene Karsch‐Mizrachi, Eric P. Nawrocki · 发表于:BMC Bioinformatics · 年份:2020 · DOI:10.1186/s12859-020-3537-3 · 被引用次数:81 · 研究领域:Influenza Virus Research Studies、Biomedical Text Mining and Ontologies、Genomics and Phylogenetic Studies

BACKGROUND: GenBank contains over 3 million viral sequences. The National Center for Biotechnology Information (NCBI) previously made available a tool for validating and annotating influenza virus sequences that is used to check submissions to GenBank. Before this project, there was no analogous tool in use for non-influenza viral sequence submissions. RESULTS: We developed a system called VADR (Viral Annotation DefineR) that validates and annotates viral sequences in GenBank submissions. The annotation system is based on the analysis of the input nucleotide sequence using models built from curated RefSeqs. Hidden Markov models are used to classify sequences by determining the RefSeq they are most similar to, and feature annotation from the RefSeq is mapped based on a nucleotide alignment of the full sequence to a covariance model. Predicted proteins encoded by the sequence are validated with nucleotide-to-protein alignments using BLAST. The system identifies 43 types of "alerts" that (unlike the previous BLAST-based system) provide deterministic and rigorous feedback to researchers who submit sequences with unexpected characteristics. VADR has been integrated into GenBank's submission processing pipeline allowing for viral submissions passing all tests to be accepted and annotated automatically, without the need for any human (GenBank indexer) intervention. Unlike the previous submission-checking system, VADR is freely available (https://github.com/nawrockie/vadr) for local ...