Estimating the relative frequencies of linguistic features is a fundamental task in linguistic computation. As the amount of text or speech that is available from a given user of the language typically varies greatly, and the sample sizes tend to be small, the most straightforward methods do not always give the most informative answers. Bootstrap and Bayesian methods provide techniques for handling the uncertainty in small samples. We describe these techniques for estimating frequencies from small samples, and show how they can be applied to the study of linguistic change. As a test case, we use the introduction of the pronoun you as subject in the data provided by the Corpus of Early English Correspondence (c. 1410–1681).
No takes yet. Share an insight, caveat, or question.
Hinneburg et al. (2006) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: