Scholay

学术搜索 · AI 审稿 · LaTeX 协作

First and second moment of counts of words in random texts generated by Markov chains

作者:J. Kleffe, Mark Yu. Borodovsky · 发表于:Computer applications in the biosciences · 年份:1992 · DOI:10.1093/bioinformatics/8.5.433 · 被引用次数:58 · 研究领域:RNA and protein synthesis mechanisms、Fractal and DNA sequence analysis、Machine Learning in Bioinformatics

An exact expression for the variance of random frequency that a given word has in text generated by a Markov chain is presented. The result is applied to periodic Markov chains, which describe the protein-coding DNA sequences better than simple Markov chains. A new solution to the problem of word overlap is proposed. It was found that the expected frequency and overlapping properties determine most of the variance. The expectation and variance of counts for triplets are compared with experimental counts in Escherichia coli coding sequences.