Comparison of the pattern of synonymous nucleotide substitution between two complete genomes of Mycobacterium tuberculosis at 3,298 putatively orthologous loci showed a mean percent difference per synonymous site of 0.000328 ± 0.000022.Although 80.5% of loci showed no synonymous or nonsynonymous nucleotide differences, the level of polymorphism observed at other loci was greater than suggested by previous studies of a small number of loci.This level of nucleotide difference leads to the conservative estimate that the common ancestor of these two genotypes occurred approximately 35,000 ago, which is twice as high as some recent estimates of the time of origin of this species.Our results suggest that a large number of loci should be examined for an accurate assessment of the level of nucleotide diversity in natural populations of pathogenic microorganisms.urveys of genetic diversity in the pathogenic bacterium Mycobacterium tuberculosis have revealed a contradictory picture.In spite of known polymorphism at the phenotypic level and abundant polymorphism associated with repetitive elements (1), surveys of single nucleotide polymorphism in protein-coding genes have shown surprisingly low levels of polymorphism in comparison with other eubacterial species (2).The apparent low level of nucleotide polymorphism has led to the hypothesis that the ancestor of this species occurred quite recently, perhaps 15,000-20,000 years ago (2,3).However, if the number of substitutions per site is low, the error of estimation of this number would be expected to be substantial unless a very large number of sites are surveyed.We addressed the question of polymorphism in M. tuberculosis by comparing protein-coding genes in two completely sequenced genotypes, H37Rv and CDC1551 (4). MethodsWe applied the BLASTP program (5) to identify, for each predicted protein sequence in the H37Rv genome (GenBank accession no.AL123456), the closest homolog in the CDC1551 genome (GenBank accession no.AE000516).Following GenBank annotations, we compared 3,972 predicted proteins in H37Rv with 4,187 predicted proteins in CDC1551.We used a strict search criterion (E = 10 -50 ) to identify truly orthologous gene pairs.We aligned (6) the putative orthologous pairs of amino acid sequences (n=3,428), then imposed this alignment on the DNA sequences.Visual inspection of amino acid alignments showed that certain alignments, usually near the N-terminus or C-terminus, had regions of very low sequence identity.Examination of the DNA sequences of the corresponding genes showed that these regions of low identity were typically caused by a frameshift in one of the two genomes relative to the other.Whether these frameshifts are biologically real or result from sequencing error was uncertain; therefore, we eliminated 119 such gene pairs from our data set.For the remaining gene pairs (n=3,309), we computed the proportion of synonymous substitutions per synonymous site (p S ) and the proportion of nonsynonymous substitutions per site (p N ) by using Nei and Gojobori's method (7).Because values of p S and p N were very low in most cases, we did not correct for multiple hits.Because p S values appeared to fall into two groups (see Results), we used a simple probabilistic model to separate these two sets of gene pairs.We assumed that the probability of synonymous substitution followed two separate binomial distributions, designated models A and B, with probabilities of "success" (i.e., of a synonymous difference) designated p A and p B , respectively.Using the Bayes equation, for each gene pair with a given p S value, we computed the probability that model A applies, given the observed p S : P(A|p S ) = (p SA ) f A /[(p SA ) f A + (p SB ) f B ], where f A is the frequency of cases to which model A applies, f B the frequency of cases to which model B applies, p SA is the binomial probability of obtaining the observed p S , given the number of synonymous sites in the gene and a probability of a synonymous difference equal to p A ; and p SB is the binomial probability of obtaining the observed p S , given the number of synonymous sites in the gene and a probability of a synonymous difference equal to p B .The probability that model B applies, givenBy comparing these two probabilities, we assigned each gene pair to one of two groups assumed to evolve according to the two models, respectively (Groups A and B).We reassigned group membership in iterative fashion, computing p A and p B from the mean p S values for each group.We started the process with f A = 0.995 and continued until group memberships were stable. ResultsOf 3,309 pairs of putatively homologous protein-coding genes in the H37Rv and CDC1551 genomes of M. tuberculosis, 2,662 (80.5%) showed no synonymous or nonsynonymous
No takes yet. Share an insight, caveat, or question.
Hughes et al. (2002) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: