Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Utilizing sequence intrinsic composition to classify protein-coding and long non-coding transcripts

作者:Liang Dan Sun, Haitao Luo, Dechao Bu, Guoguang Zhao, Kuntao Yu, Changhai Zhang, Yuanning Liu, Runsheng Chen, Yi Chen Zhao · 发表于:Nucleic Acids Research · 年份:2013 · DOI:10.1093/nar/gkt646 · 被引用次数:2380 · 研究领域:Genomics and Phylogenetic Studies、Cancer-related molecular mechanisms research、Genetic diversity and population structure

It is a challenge to classify protein-coding or non-coding transcripts, especially those re-constructed from high-throughput sequencing data of poorly annotated species. This study developed and evaluated a powerful signature tool, Coding-Non-Coding Index (CNCI), by profiling adjoining nucleotide triplets to effectively distinguish protein-coding and non-coding sequences independent of known annotations. CNCI is effective for classifying incomplete transcripts and sense-antisense pairs. The implementation of CNCI offered highly accurate classification of transcripts assembled from whole-transcriptome sequencing data in a cross-species manner, that demonstrated gene evolutionary divergence between vertebrates, and invertebrates, or between plants, and provided a long non-coding RNA catalog of orangutan. CNCI software is available at http://www.bioinfo.org/software/cnci.