DNA sequences of higher organisms contain thousands of nearly identical dispersed repetitive sequences. In order to understand the effect of such repeats on word entropies, we construct a model that can be analyzed analytically. The hypothetical model sequences consist of independent equidistributed symbols with randomly interspersed repeats. As a conclusion, we predict that the entropy of DNA sequences measuring the information content is much lower than suggested by earlier empirical studies.
No takes yet. Share an insight, caveat, or question.
Herzel et al. (1994) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: