Sir, The mechanisms of DNA replication initiation are quite different in bacteria and eukarya. In bacteria, initiation occurs at a single locus, oriC, and is triggered by a single protein, DnaA (Kornberg and Baker, 1992, DNA Replication. New York: W. H. Freeman and Co.), whereas, in eukarya, initiation takes place at multiple replication origins that are permanently occupied by origin of replication complexes (ORC) made up of five or six protein subunits. These ORCs are made competent for initiation by the loading of minichromosome maintenance proteins (MCM) (Kearsey and Labib, 1998, Biochim Biophys Acta1398: 113–136; Pasero and Gasser, 1998, Curr Opin Cell Biol10: 304–310), an association that is triggered by the protein Cdc6 (Liang and Stillman, 1997, Genes Dev11: 3375–3386). Until recently, nothing was known about the initiation of DNA replication in the third domain of life, the Archaea, not even whether they have single or multiple replication origins (Edgell and Doolittle, 1997, Cell89: 995–998). This situation is changing with the advent of archaeal genomics. In particular, Cdc6/Orc1 and MCM homologues have been detected in completely sequenced archaeal genomes (Bernander, 1998, Mol Microbiol4: 955–961). In silico attempts have recently been made to identify replication origins in archaeal chromosomes (Grigoriev, 1998, Nucleic Acids Res26: 2286–2290; Salzberg et al., 1998, Gene217: 57–67). First, in some bacteria, the leading strand contains more G than C, so that the origin (oriC) and terminus (terC) of chromosome replication can be detected by plotting this GC skew along the genome (Lobry, 1996, Mol Biol Evol13: 660–665). Grigoriev (1998, Nucleic Acids Res26: 2286–2290) improved this method by the use of cumulative diagrams that display two peaks when there is a unique origin of replication. Among Archaea, cumulative GC skew diagrams suggested a single oriC only for Methanococcus jannaschii and Methanobacterium thermoautotrophicum. In M. thermoautotrophicum, Grigoriev noticed a homologue of the bacterial chromosome partition gene soj close to one of the two peaks, but failed to identify chromosome regions bearing consensus sequences for potential replication origins. Second, some oligomers also appear to have a skewed distribution along the genome. Salzberg et al. (1998, Gene217: 57–67) have located the origin of replication by maximizing the overall oligomer skew on half genomes. In M. thermoautotrophicum, Salzberg and co-workers identified the same region of potential replication origin as Grigoriev and noticed the presence of an archaeal homologue of the eukaryotic DNA replication initiator gene cdc6/orc1 at 5 kb from this putative origin, but they failed to identify skewed oligomers in other Archaea. We have applied the cumulative skew technique to oligomers from two to eight nucleotides (words) for all completely sequenced prokaryotic genomes. Cumulative word skews give more precise diagrams than GC skews (Fig. 1). An explanation could be that when genome rearrangements occur the best word skew tends to be restored more rapidly than the GC skew because the latter is probably due to a mutational bias, which is slow because it is selectively neutral. In contrast, the best word skew can be related to the preferential location on the leading strand of signals that are under selective pressure, such as primase recognition sites (Blattner et al., 1997, Science277: 1453–1462). This technique allowed us to find the origin of replication even when the cumulative GC skew was not informative (Aquifex aeolicus, data not shown). . Cumulative word, GC and codon skew diagrams for Methanobacterium thermoautotrophicum and Pyrococcus horikoshii. For every word (W), non-overlapping occurrences of W and of its reverse complement Wt (e.g. AATCG for CGATT) were counted on a sliding window along the genome. The location of the window was incremented by 1/240th of the genome for better precision, and its size was 1/50th of the genome for better smoothing, conditions similar to those described previously by Grigoriev (1998, Nucleic Acids Res26: 2286–2290). Cumulative skew diagrams were obtained by integrating (nW − nWt)/(nW + nWt) over the 240 locations, in which nW is the number of occurrences of W. Best words were then selected on the smoothness of their cumulative diagram. Our analyses have shown that words of four nucleotides were sufficient for conclusive results. Similarly, W is assumed to be nucleotide G for GC skew (Wt is C). For codon skew, nW is the number of codons in the sliding window that are transcribed in the arbitrary positive sense. Ordinates are arbitrary units (cumulated skews), abscissas represent the position in the genome (start is given by complete genome sequences). The arrows indicate the position of Cdc6/Orc1 homologues MT1412 and PH0124 in the complete genome sequences. Genomes were obtained from Smith et al. (1997, J Bacteriol179: 7135–7155) and Kawarabayasi et al. (1998, DNA Res5: 55–76). We obtained two-peaked diagrams for the archaea Pyrococcus horikoshii (best word GGGT) and M. thermoautotrophicum (best word GGCA) (Fig. 1). In the latter case, the peaks were slightly different from those previously found by Grigoriev (1998, Nucleic Acids Res26: 2286–2290). Because cumulative word skews clearly divide the genomes of M. thermoautotrophicum and P. horikoshii into two halves, these peaks were promising candidates for origin/termination of bidirectional replication. However, neither GC nor cumulative word skews allow determination of origins for the archaea M. jannaschii or Archaeoglobus fulgidus, or for the bacterium Synechocystis. Considering that initiator genes are very often located close to the origin in bacteria, plasmids and viruses, we looked at the ORFs located in the regions surrounding the peaks. Most interestingly, these regions contained archaeal homologues of cdc6/orc1 in M. thermoautotrophicum (accession number MT1412) and in P. horikoshii (accession number PH0124) (Fig. 1). In contrast, the soj gene previously noticed by Grigoriev (1998, Nucleic Acids Res26: 2286–2290) is located 40 kb and 550 kb away from the cdc6/orc1-containing peaks of M. thermoautotrophicum and P. horikoshii respectively. As observed at the origin of replication in most bacteria, the plot of G − C/G + C shifted from negative to positive at the peak close to the cdc6/orc1 gene, in agreement with this location being the origin. Moreover, all rRNA genes and the majority of ribosomal protein genes were transcribed in the same direction as DNA replication under our working hypothesis (not shown). Such organization is expected to avoid head-to-head collisions between transcription complexes and replication forks (French, 1992, Science258: 1362–1365). Indeed, using cumulative diagrams again, the codon skew correlated reasonably well with the cumulative word skew both in M. thermoautotrophicum and P. horikoshii (Fig. 1). The congruence between the location of a putative initiator gene and GC, oligomer and codon skews prompted us to look for putative origin sequences surrounding the cdc6/orc1 gene. Replication origins are usually located in large intergenic regions that exhibit complex patterns of AT-rich elements as well as direct and inverted repeats (Kornberg and Baker, 1992, DNA Replication, New York: W. H. Freeman and Co.). Indeed, we found such regions at the 5′ end of the two archaeal cdc6/orc1 genes identified by the word skew diagrams (Fig. 2). In M. thermoautotrophicum, this region contains 12 copies of a 13 bp repeat that are symmetrically distributed around an internal AT-rich element (Fig. 2). The spacing of these repeats is remarkably regular (about 50 bp), and five of them are included in larger perfect repeats of 25 and 21 bp respectively. The corresponding region of P. horikoshii also contains repeats of 13 bp distributed on either side of two central AT-rich elements. These Pyrococcus repeats turned out to be strikingly similar to those detected in M. thermoautotrophicum (Fig. 2). As expected for essential regulatory elements, the repeats and AT-rich elements of the putative P. horikoshii oriC were conserved in Pyrococcus furiosus (Fig. 2 and data not shown). . Schematic physical maps of oriC and alignments of the repeats. A. Repeats similar in the three oriC are represented by black arrows whose direction indicates sense. Longer direct repeats are present in M. thermoautotrophicum and pFZ1 oriC, but have not been represented for simplicity. Grey boxes represent central AT-rich elements whose sequences are listed. B. Repeats are aligned for the three oriC we identified, and displayed with their number of occurrences. A consensus sequence is shown. M. t., Methanobacterium thermoautotrophicum; pFZ1, plasmid pFZ1 from Methanobacterium thermoformicicum; P. h., Pyrococcus horikoshii; P. f., Pyrococcus furiosus. Accession number for pFZ1 is GenBank X67212. Interestingly, an archaeal homologue of cdc6/orc1 has been previously detected in the plasmid pFZ1 from Methanobacterium thermoformicicum (ORF1, SWISSPROT P29570) (Edgell and Doolittle, 1997, Cell89: 995–998). We checked whether repeats similar to those identified in the putative archaeal oriC were also present in pFZ1. Indeed we found seven copies of a 14 bp repeat that matches very well with archaeal oriC consensus sequences (Fig. 2). They are again present in an intergenic AT-rich region at the 5′ end of the cdc6/orc1 gene. This suggested that an archaeal chromosomal replication origin is used by pFZ1 for its own replication. The presence of both these sequences and a cdc6/orc1 homologue on a plasmid suggests that Cdc6/Orc1 itself recognizes the repeated sequences found in M. thermoautotrophicum, Pyrococcus and pFZ1 origins. The latter notion can be further supported by recent observations indicating that ORC1, ORC4 and ORC5 subunits of the ORC complex all appear to be related to each other (Tugal et al., 1998, J Biol Chem273: 32421–32429), and that all of these subunits are known to interact with ARS sequences in yeast (Lee and Bell, 1997, Mol Cell Biol17: 7159–7168). Therefore, the function(s) of archaeal Cdc6/Orc1 homologues could be more similar to that of Orc1, although they are not more similar in sequence to Orc1 than to Cdc6 eukaryotic proteins. In P. furiosus, the cdc6/orc1 gene is co-transcribed with the genes encoding the subunits DP1 and DP2 of a recently identified archaeal specific DNA polymerase (Uemori et al., 1997, Genes Cells2: 499–512). The promoter region thus defined overlaps with the oriC sequences identified in the two Pyrococcus (not shown), suggesting an interplay between control of DNA replication initiation and transcription of initiator and replicator genes. Interestingly, the M. thermoautotrophicum DP1 gene (accession number MT1405) is located only 5 kb away from the putative oriC. This could facilitate the formation of the replication complex at the origin by promoting a direct interaction between archaeal Cdc6/Orc1 and DP1. Altogether, our results strongly suggest that we have identified the unique oriC of M. thermoautotrophicum, P. horikoshii and P. furiosus, as well as consensus sequences recognized by initiator protein(s). Considering the conservation of these features between Pyrococcus species and M. thermoautotrophicum, it might seem surprising that we failed to identify putative oriC in A. fulgidus and M. jannaschii. One possible explanation is that these archaea use multiple origins instead of a single oriC. However, we suppose that all prokaryotes have a single origin of replication but that the in silico approach can be confused by very frequent genome rearrangements and/or by intrinsic properties of the genome, such as mutational bias or primase recognition site (to be detailed elsewhere). Surprisingly, the genome of M. jannaschii does not contain a Cdc6/Orc1 homologue, but four homologues of MCM instead of only one in other Archaea (Bernander, 1998, Mol Microbiol4: 955–961). This suggests that the apparatus for replication initiation is more flexible than often expected and can be affected by non-orthologous displacement. In agreement with this conclusion, it has been recently shown than disruption of the dnaA gene Synechocystis has no phenotypic effect (Richter et al., 1998, J Bacteriol180: 4946–4949), indicating that another protein must be the replication initiator in this bacterium. This opens new perspectives for studying archaea. For example, archaeal minichromosomes resembling pFZ1 could be the starting point for the construction of cloning vectors. They might be used to analyse the mechanisms of DNA replication initiation, elongation and cell cycle regulation in vitro. Such studies may have important consequences for our understanding of these processes in eukaryotes. This work was supported by research grants from the European Union Biotechnology programme (BIO4-CT 96-0488) and by the Association de Recherche contre le Cancer. H.M. is supported by a Marie Curie postdoctoral training grant from the European Union.
No takes yet. Share an insight, caveat, or question.
Lopez et al. (1999) studied this question.