The HMG-box is an approximately 75-amino acid residue protein domain that occurs in all eukaryotic organisms and was first identified as a characteristic feature of the chromosomal high mobility group (HMG) proteins of the HMGB type. Structural studies have demonstrated that the L-shaped fold of the domain formed by three α-helices is conserved to a greater extent than expected from amino acid sequence similarity between different HMG-boxes. The long arm consists of helix III and the N-terminal extended strand, whereas the short arm of the L-shape is composed of helices I and II forming an angle of approximately 80° between the arms (Thomas and Travers, 2001; Stros et al., 2007). The HMG-box domain mediates DNA binding primarily through the minor groove of DNA. Hydrophobic residues of the concave face of the L-shaped molecule partially intercalate between the DNA bases, thereby widening the minor groove, which results in unwinding and remarkable bending of the DNA helix. Thus, the HMG-box domain binds the outside of the DNA bend, compressing the major groove (Thomas and Travers, 2001; Stros, 2010). Some HMG-box proteins can interact with DNA sequence specifically (e.g. mammalian transcription factors such as SEX DETERMINING REGION OF Y [SRY] and LYMPHOID ENHANCHER-BINDING FACTOR1 [LEF-1]), whereas other HMG-box proteins bind DNA sequence independently (e.g. chromosomal HMGB proteins and Structure-Specific Recognition Protein1 [SSRP1]). A typical feature of both types of HMG- box domains is their selective binding to certain DNA structures, including four-way junctions and DNA minicircles (Bustin, 1999; Thomas and Travers, 2001; Stros et al., 2007; Wegner, 2010). Because HMG-box proteins induce DNA bending upon binding to linear DNA, they often act as architectural facilitators in the assembly of nucleoprotein complexes involved in transcription, recombination, or other DNA-dependent processes (Bustin, 1999; Thomas and Travers, 2001; Stros et al., 2007). HMG-box domains are found in a variety of proteins that interact with DNA. In these proteins, the HMG-box domain(s) occurs in combination with various other protein domains of different function. Accordingly, because of this structural variability and their interaction with various other proteins, HMG-box proteins are involved in different nuclear functions. There are HMG-box proteins, for instance, that act as architectural chromosomal proteins (HMGB proteins), whereas others are transcription factors or subunits of chromatin-remodeling complexes, or they modulate DNA recombination/repair (Bustin, 1999; Stros et al., 2007). In addition to the cell nucleus, HMG-box proteins are found in mitochondria of animals and yeast, where they serve as transcriptional regulators and contribute to the organization of the mitochondrial DNA (Bonawitz et al., 2006; Kucej and Butow, 2007). Currently, there is no evidence for the occurrence of HMG-box proteins in plant mitochondria. However, an unusual HMG-box protein from Physcomitrella localizes to plastids in tobacco (Nicotiana tabacum) BY-2 cell protoplasts (Kiilerich et al., 2008), but no higher plant HMG-box protein has been reported to occur in plastids. Various plant genomes were found to encode HMG-box proteins, suggesting that they commonly occur in plants (Riechmann et al., 2000; Stros et al., 2007). Higher plant genomes encode 10 to 15 different HMG-box proteins that range from approximately 13 to 72 kD. When compared with the human genome, which encodes 47 HMG-box proteins ranging from approximately 15 to 193 kD, HMG-box proteins are less diversified in plants (Stros et al., 2007). Whereas in humans, HMG-box-containing transcription factors represent the largest subgroup (Stros et al., 2007; Wegner, 2010), to date, it is unclear whether any of the plant HMG-box proteins act as transcription factors. No sequence-specific DNA interactions have been reported for any of the plant HMG-box proteins. In plants, the family of small chromosomal HMGB proteins represents the most diversified subgroup of HMG-box proteins (Stros et al., 2007). We have searched various databases, including the Plant Chromatin Database (www.chromdb.org/) and that of MIPS (database from the Munich Information Center for Protein Sequences; http://mips.helmholtz-muenchen.de/plant/), for plant proteins containing the HMG-box motif. In addition to flowering plants, we investigated the extent to which different types of HMG-box proteins are encoded in the genome sequences of the lycophyte Selaginella moellendorfii, the moss Physcomitrella patens, and the algae Chlamydomonas reinhardtii and Volvox carteri that became available in recent years. The filtered results of these searches were used to compile a relatively comprehensive list of plant HMG-box proteins (Supplemental Table S1). Based on their overall structure and amino acid sequence similarity (Fig. 1), plant HMG-box proteins can be subdivided into four distinct families: chromosomal HMGB proteins, AT-rich interaction domain (ARID)-HMG proteins, 3xHMG-box proteins, and SSRP1. We used the amino acid sequences of HMG-box proteins from nine species (three monocots, three dicots, Selaginella, Physcomitrella, and Chlamydomonas) for a multiple sequence alignment that served for the construction of a neighbor-joining tree (Fig. 1). It clearly illustrates the four distinct families of plant HMG-box proteins. Various studies performed in the past few years suggest that members of the four HMG-box protein families have different cellular functions, and we discuss here the present knowledge about the structural and functional characteristics of these proteins. Schematic representation of the overall structure of the four families of plant HMG-box proteins and their amino acid sequence similarity. Whereas HMGB proteins, ARID-HMG proteins, and SSRP1 contain a single HMG-box domain, the 3xHMG-box proteins have three copies of the HMG-box domain. The overall domain structure of the four groups of HMG-box proteins that were identified in plants are presented schematically: HMG-box domain (blue), basic region (green), acidic region (red), SSR domain of SSRP1 (orange), and ARID (violet). The amino acid sequences of HMG-box proteins from Brachypodium distachyon (Bd), rice (Os), maize (Zm), Arabidopsis (At), P. trichocarpa (Pt), grape (Vv), S. moellendorfii (Sm), P. patens (Pp), and C. reinhardtii (Cr; compare with Supplemental Table S1) were aligned by multiple sequence alignment that served for the construction of a neighbor-joining tree using the software package SeaView (http://pbil.univ-lyon1.fr/software/seaview.html). The four families of plant HMG-box proteins, HMGB (in black), ARID-HMG (in blue), 3xHMG-box (in green), and SSRP1 (in red), occur as distinct groups. Whereas the proteins of Selaginella and Physcomitrella group with their counterparts from flowering plants, the two Chlamydomonas HMGB-type sequences (in violet) group separately. Originally, HMG proteins were identified as proteins with unusual physicochemical properties when calf thymus chromatin was extracted with 0.35 m NaCl (Goodwin et al., 1973). Subsequently, based on their characteristic amino acid sequences, they were subdivided into three structurally distinct families termed HMGA, HMGB, and HMGN (Bustin and Reeves, 1996; Grasser et al., 2007a). In this article, we concentrate exclusively on the HMG-box containing HMGB family. The HMGB proteins (13–27 kD) of different organisms exhibit a diverse overall structure. The vertebrate proteins, for instance, consist of two HMG-box domains, a basic linker region and an acidic C-terminal domain, whereas plant HMGB proteins (Fig. 1) contain a single HMG-box domain that is flanked by basic N-terminal and acidic C-terminal domains (Thomas and Travers, 2001; Stros et al., 2007). Database analyses revealed that HMGB-type proteins apparently occur in all plants and also in algae (Supplemental Table S1). Plant HMGB proteins are structurally more diversified than their animal counterparts (Stros et al., 2007). Thus, the Arabidopsis (Arabidopsis thaliana) genome encodes eight proteins that, according to their amino acid sequences, can be classified as HMGB proteins. However, experimental analyses demonstrate that the protein encoded by the Arabidopsis Genome Initiative locus At5g23405, despite marked sequence similarity to well-characterized HMGB proteins, does not share the features of bona fide HMGB proteins, including predominant nuclear localization and DNA-binding activity (Grasser et al., 2006). Therefore, it is required to test experimentally the functionality of predicted HMG-box domains, which may be of particular importance for domains that display a lower degree of sequence conservation such as the putative algal HMG-box proteins. Plant HMGB proteins and their counterparts from other sources essentially have in common the characteristic DNA interactions such as low affinity, sequence-independent binding to linear DNA, but high-affinity interaction with DNA structures (four-way junctions, DNA minicircles, supercoiled DNA) and pronounced DNA bending upon binding linear DNA. We have previously reviewed these aspects of DNA binding (Grasser et al., 2007a) and pointed out that various HMGB proteins (e.g. those occurring in maize [Zea mays] and Arabidopsis) display differences in their DNA interactions and posttranslational modifications. These differences indicate that plants have a repertoire of architectural chromatin-associated proteins. DNA-binding experiments with full-length and truncated proteins also indicated that interactions between the N-terminal basic domain (which increases DNA binding) and the acidic C-terminal domain (which reduces DNA binding) regulate the DNA interactions of maize and rice (Oryza sativa) HMGB proteins (Ritt et al., 1998; Wu et al., 2003). Spectrometric measurements and cross-linking experiments confirmed intramolecular interactions between the terminal domains of plant HMGB proteins and that the interactions are enhanced by the phosphorylation of Ser residues within the acidic tail (Thomsen et al., 2004). Consistent with these findings, removal of the acidic tail of Arabidopsis HMGB1 and HMGB5 reduced the remarkable mobility of the proteins within cell nuclei in living cells as measured by fluorescence recovery after photobleaching (FRAP) experiments (Launholt et al., 2006). The intramolecular domain interactions appear to influence various properties of plant HMGB proteins, including DNA interactions, mobility within the nucleus, and subcellular localization (Ritt et al., 1998; Wu et al., 2003; Launholt et al., 2006; Pedersen et al., 2010). Although it is well documented that mammalian HMGB1 can be detected also outside the nucleus, acting as a kind of cytokine (Müller et al., 2004; Yang and Tracey, 2010), chromosomal HMGB proteins are generally considered nuclear proteins (Grasser et al., 2007a; Reeves, 2010). In line with that, plant HMG proteins traditionally were purified from chromatin or isolated nuclei (Spiker, 1984; Grasser et al., 1991). Systematic examination of the subcellular localization of Arabidopsis HMGB1, HMGB5, and HMGB6 proteins as well as of the HMGB-type protein AtHMGB14 by analyzing the distribution of GFP fusions and by immunofluorescence microscopy confirmed the nuclear localization (Grasser et al., 2004, 2006; Launholt et al., 2006; Lildballe et al., 2008). In contrast to these findings, recent experiments revealed that Arabidopsis HMGB2/3 and HMGB4, in addition to being found in the nucleus, are detected to different extents in the cytosol. Monitoring the distribution of photoactivatable GFP fused to HMGB2 and HMGB4 demonstrated that both proteins can shuttle between nucleus and cytoplasm, whereas HMGB1 remained nucleus localized (Pedersen et al., 2010). Currently, it is unclear why some plant HMGB proteins are strictly nuclear whereas others can shuttle between nucleus and cytosol. An extranuclear/extracellular role like that of HMGB1, which acts as a specific mediator in injury and inflammation of mammals (Müller et al., 2004; Yang and Tracey, 2010), appears unlikely for plants. In view of the interplay between HMGB proteins and linker histones (see below), the nucleocytosolic of HMGB proteins may serve as a to regulate the nuclear of these architectural proteins and thereby modulate the with linker histones that may influence chromatin structure. the cell nucleus, HMGB proteins are proteins that bind on to the binding and and 2010). Whereas linker increases chromatin HMGB other HMG that share with bind chromatin more than chromatin the of factors to chromatin (Bustin et al., and Thomas and the interplay between and HMGB proteins, et proteins (e.g. into cells the other protein as a GFP (e.g. The mobility to that chromatin of the GFP was measured by in cells and these experiments revealed that in living mammalian HMGB proteins other HMG with linker for chromatin In line with that, mammalian HMGB1 in with linker histones acidic C-terminal domain, which may in an enhanced DNA binding of HMGB1 et al., 2008). In the chromatin this the of by HMGB proteins, in a more chromatin structure that is for transcription (Bustin et al., and Thomas and by Arabidopsis HMGB1 and HMGB5 are nuclear proteins with a high on chromatin and a short is clearly higher than that of linker HMGB proteins the acidic C-terminal domain in these experiments a reduced suggesting that the for DNA of the truncated proteins their chromatin interactions (Launholt et al., 2006). Therefore, in plant cell there is a range of HMGB proteins that, in with other chromosomal proteins, can modulate chromatin structure. In addition to the interplay with linker HMGB proteins as are involved in the assembly of specific nucleoprotein complexes such as transcription and 2003; Grasser et al., 2007a). The of plant HMGB proteins to the of specific nucleoprotein structures is from their architectural role in which was in and in et al., HMGB proteins can the binding of certain plant transcription factors DNA binding with to their DNA by interaction (in their with these factors et al., 1996; In the of maize the of the interaction is by protein phosphorylation and is by the terminal domains of the HMGB proteins et al., Grasser et al., HMGB5 (which represents the most of DNA binding) can of the transcription also in the et al., 2003). with HMGB proteins of different sources have indicated that they are involved as architectural factors in various DNA-dependent nuclear including transcription, recombination, and DNA to their and interaction with DNA and other proteins, HMGB proteins have the to these processes by chromatin structure by the of nucleoprotein complexes Stros, 2010). In plants, it has that of an role in the to and it appears that linker histones and HMGB proteins are involved in the et al., 2010). The of some Arabidopsis HMGB is by et al., 2007). Arabidopsis plants that Arabidopsis HMGB2 a HMGB upon different reduced whereas of HMGB4 no marked The (in may be to the of a of et al., 2007; et al., 2008). the and of HMGB1 in Arabidopsis in an When to NaCl the of HMGB1 was whereas the of plants compared with plants revealed that a of were the the of was et al., 2008). of plants to a in the Chromatin HMGB proteins, linker and other chromatin may plant in to the different types of et al., 2010). In a role of HMGB proteins in and was of (which is in the of cells to and Some of the in the cells were found to encode putative plant of the that in mammals is involved in et al., The of this to be role of plant HMGB proteins was identified in the of Arabidopsis HMGB1 whereas in plants HMGB1, were In contrast to cells HMGB1, in of the apparently is not by reduced activity et al., Currently, the of of HMGB1 in Arabidopsis is but it is that HMGB proteins other than HMGB1 influence the chromatin structure. to the of HMGB proteins in plant cell group of HMG-box proteins that was identified by sequence searches are the proteins containing in addition to an HMG-box a ARID (Riechmann et al., 2000; Stros et al., 2007). The ARID is a DNA-binding that was first identified in the transcription and the were found in a variety of animal transcription factors that regulate cell and The ARID sequence approximately amino acid residues that are in a fold et al., 2000; et al., bind DNA through a that major groove of a which the helices of the and through structural from a The ARID proteins were found to bind to AT-rich sequences, the of the domain. of interactions of other by various revealed that of different proteins bind DNA specifically et al., et al., Plant genomes encode various ARID proteins, and in most the ARID occurs in combination with other protein domains, including and HMG-box domains (Riechmann et al., of the ARID protein from that an ARID no other demonstrated that it binds AT-rich of the It has been that a role in to the plant et al., 2008). The combination of an N-terminal ARID and a C-terminal HMG-box in the ARID-HMG proteins kD, a Physcomitrella protein with 1) is specific for plants, occurring in flowering plants as well as Selaginella and Physcomitrella, but apparently not in algae with Supplemental Table S1). ARID-HMG proteins to be more diversified in species than in the amino acid sequences of ARID-HMG proteins from various species indicated the of distinct structural (Supplemental S1). of the four Arabidopsis ARID-HMG termed were found to be in the different In BY-2 and primarily to the a binding to AT-rich DNA compared with and by HMG-box domain it can bind DNA structure the ARID and the HMG-box domain contribute to the DNA interactions of et al., 2008). Because both the ARID and the HMG-box occur in a variety of proteins with different functions, the role of the ARID-HMG proteins in plants is and The 3xHMG-box proteins kD) are composed of an N-terminal basic domain (which no similarity to other and three copies of the HMG-box domain in the C-terminal (Fig. 1). to the 3xHMG-box proteins occur exclusively in plants, but they are relatively conserved plants, because they are encoded in genomes of lower plants and higher plants with Supplemental Table S1). Whereas some species encode two proteins (e.g. Arabidopsis and other species encode a single 3xHMG-box protein (e.g. rice and grape We were to 3xHMG-box sequences from but Chlamydomonas apparently encodes proteins with two putative HMG-box domains, which not occur in plants. In the of experimental it is whether these proteins to the HMGB family HMGB proteins have two or whether they represent the of the 3xHMG-box proteins. We have used the amino acid sequences of the HMG-box domains of the 3xHMG-box proteins with the sequences of the HMG-box domains of HMGB, and ARID-HMG proteins to a neighbor-joining revealed that the various HMG-box sequences, despite their overall group strictly according to the protein family they from (Supplemental the different HMG-box sequences and of 3xHMG-box proteins are clearly forming the two Arabidopsis 3xHMG-box proteins, termed and which share amino acid sequence were The 3xHMG-box proteins are in the but in contrast to other HMG-box proteins, they are in a in of GFP proteins and immunofluorescence studies demonstrated that and with different of (Pedersen et al., clearly the 3xHMG-box proteins from other HMG-box proteins (e.g. Arabidopsis HMGB proteins and that interact with chromatin but not with et al., 2004; Launholt et al., 2006; Lildballe et al., Pedersen et al., 2010). generally binds the whereas with specific chromosomal the DNA (Fig. In addition to the 3xHMG-box proteins are in cells and interact with in cells (Pedersen et al., The of the 3xHMG-box proteins with both and for a role in to cell such as The of chromatin into is of importance for et al., In line with a role in the three HMG-box domains as well as the basic N-terminal domain contribute to the DNA interactions of (Pedersen et al., Therefore, the 3xHMG-box protein can interact with DNA in combination with other may contribute to the of DNA and chromatin it was and that 3xHMG-box proteins may with linker as for HMGB proteins in the binding of in a feature that was for plant However, there are other and that to be such as the of why some plants have two 3xHMG-box proteins and other plants have a single Arabidopsis 3xHMG-box proteins with cells of Arabidopsis plants or the of the were by immunofluorescence using an and for a and the of the two is Whereas generally with the of a cell is detected chromosomal by in were identified as DNA Pedersen et al., In cells an also 3xHMG-box proteins to the of a cell but no is in cells SSRP1 with the protein the chromatin transcription that was identified in and mammalian cells et al., 1998; et al., 1999; et al., The the that represent to transcription and other DNA-dependent It acts as a that of chromatin by in the of the is also involved in of thereby the chromatin from within the from various that these chromatin are by in with other complexes, and and 2006; and for SSRP1 are conserved in flowering plants as well as Selaginella and Physcomitrella and occur also in algae with Supplemental Table S1). Some plant species (e.g. Arabidopsis and have a single whereas others and P. have two proteins amino acid sequence Plant SSRP1 kD) has an overall structure (Fig. 1) to that of animal with sequence similarity the and acidic domains as well as the C-terminal HMG-box domain, but the approximately amino acid residues that the C-terminal region of animal SSRP1 et al., in yeast, the is as an HMG-box domain, is by a small HMGB-type protein termed and A short basic region terminal to the HMG-box domain is for the nuclear localization of maize SSRP1. SSRP1 binds DNA sequence but by the HMG-box domain it the structure of supercoiled and DNA as well as and the DNA interactions are by the phosphorylation of SSRP1 by protein et al., 2000; and 2001; et al., 2003). The of the of SSRP1 and in Arabidopsis was demonstrated by of the two proteins and by with specific for SSRP1 and et al., 2004). In line with role in transcription, Arabidopsis SSRP1 is detected by immunofluorescence analyses in of the nucleus but not in chromatin was found to with the region of in a but not with or et al., 2004; and 2007). Arabidopsis SSRP1 is an et al., 2010), which the in and yeast, where of the for cell and 2004; of SSRP1 various in Arabidopsis and Thus, plants display an of and have and and their is et al., 2010). The of the is clearly reduced in and plants both and the of the (which is by is that the of the in Arabidopsis is by reduced et al., 2010). purified mammalian protein complexes in a in transcription revealed that et al., 2006). and in to the of that is for the of within the sequences of et al., 2008), but the of the interplay between and is unclear et al., 2007). In interactions were when in the of subunits and were with the single and plants. and appear to act independently in the of but they act in and the of the and were to et al., 2010). These indicated multiple of interaction between Arabidopsis and and that the two factors in the of some whereas they act independently studies of the interaction between and other chromatin may into the interplay of factors that chromatin for A role of SSRP1 that was in Arabidopsis is in SSRP1 is required for DNA and for the of in the which represents the cell of et al., The that SSRP1 may be involved in the chromatin which DNA in the cell by the DNA this appears to be of in contrast to no be in et al., the overall of and is et al., 2010). It has been reported that mammalian SSRP1 also has in that are of for instance, as a transcriptional et al., 1999; et al., et al., 2007). the of SSRP1 appear to be to the In with that, the Arabidopsis subunits were found to et al., and the of SSRP1 and the plant other (Supplemental Therefore, it be to not the role of the HMG-box protein SSRP1 as of the but also functions. containing HMG-box domain(s) occur in different and there are the HMGB proteins and which apparently occur in all whereas HMG-box transcription factors to be specific for animals and the 3xHMG-box and ARID-HMG proteins appear to be plant In addition to their structural the interactions with DNA and other proteins contribute to the multiple that HMG-box proteins in the cell HMG-box domains often act as architectural in chromatin structure the assembly of higher nucleoprotein complexes as a for DNA-dependent processes in the nucleus are or by HMG-box proteins, including the and of transcription, DNA and chromatin and DNA (Bustin, 1999; Thomas and Travers, 2001; and Stros et al., 2007; Wegner, of HMG-box proteins are also from studies in plants, such as the role of SSRP1 in et al., or the specific of 3xHMG-box proteins with (Pedersen et al., Therefore, it can be expected that in plants to a of the of of the different HMG-box proteins as well as to the of functions. The are available in the of this Supplemental acid sequence similarity of proteins. Supplemental acid sequence similarity of the HMG-box Supplemental of the the Supplemental Table Plant proteins containing HMG-box We for high mobility group fluorescence recovery after photobleaching AT-rich interaction domain chromatin transcription
No takes yet. Share an insight, caveat, or question.
Antosch et al. (2012) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: