Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Estimating an Author's Vocabulary

作者:Donald R. McNeil · 发表于:Journal of the American Statistical Association · 年份:1973 · DOI:10.1080/01621459.1973.10481342 · 被引用次数:55 · 研究领域:Bayesian Methods and Mixture Models、Authorship Attribution and Profiling、Algorithms and Data Compression

The problem of estimating an author's vocabulary, given a sample of the author's writings, is considered. It is assumed that the vocabulary is fixed and finite, and that the author writes a composition by successively drawing words from this collection, independently of the previous configuration. Attention is focussed on the random variable X(n), the total number of different words used in a sample of n. It is shown that under fairly general conditions, the distribution of X(n), suitably normalized and scaled, is asymptotically Gaussian, and this result may be used to obtain a large sample estimator of vocabulary size.