Glycosyl, or glycoside, hydrolases (GHs) comprise a structurally diverse group of enzymes that hydrolyze glycosidic bonds between carbohydrates, or between carbohydrates and other noncarbohydrate moieties, and that collectively exhibit a wide range of substrate specificities. GH enzymes from across the taxonomic spectrum were originally named based on substrate specificity, the corresponding International Union of Biochemistry and Molecular Biology (IUBMB) nomenclature system (EC 3.2.1.-), and the chronological order in which they were reported. However, as growing numbers of GH proteins, and later genes, were characterized, this strategy proved to be increasingly unsatisfactory and a complementary nomenclature was developed based on predicted protein sequence (Henrissat et al., 1998). This approach provides important insights into protein structure, evolutionary relationships, and an opportunity to infer mechanistic relationships. A regularly updated database, Carbohydrate-Active Enzymes (CAZY; www.cazy.org), currently lists 108 distinct GH families, a subset of which are further affiliated with 14 clans based on the presence of defined protein folds and conserved catalytic machineries. This nomenclature initiative was developed largely in response to the rapidly growing numbers of reported microbial GHs, reflecting their numerous important industrial uses. Notable examples are endo-β-1,4-glucanases, or cellulases, which hydrolyze the β-1,4-glucosyl linkages of cellulose and several other plant cell wall polysaccharides, and are used in the generation of sugars from lignocellulosic biomass for ethanol production. Plant biologists are currently facing a similar nomenclature dilemma, since the sequencing of whole plant genomes has revealed many large GH families (Henrissat et al., 2001): a genome-scale assessment of Arabidopsis (Arabidopsis thaliana) in CAZY identified 393 GHs from 34 families (http://www.cazy.org/geno/3702.html), and equivalent families are emerging in many species as larger EST collections develop. GH activities from plants have long been studied in association with various aspects of growth, development, and cell wall metabolism, but plant GH enzymes have often been named with little consideration of their substrate specificity, molecular structure, or enzymatic reaction mechanism. The current explosion of interest and new research opportunities in biofuels and bioenergy crops will inevitably result in renewed interest in identifying and annotating plant GHs, through their association with lignocellulosic biomass. Thus, this is an opportune time to adopt a new, rational, well-defined nomenclature. Importantly, it is apparent that many of these plant enzymes share similarities to previously classified families of microbial GHs, which is not clear from their original designations. While the IUBMB-recommended nomenclature (www.chem.qmul.ac.uk/iubmb), based on the substrates used and the reaction catalyzed, continues to provide critical information regarding enzyme function, we propose a complementary, standardized nomenclature for plant GHs, based on the rigorous scheme for naming their microbial counterparts, which is categorized by the catalytic domain of the enzyme (Henrissat et al., 1998). Schematic representation of the modular plant GH9 family structure. Specific domains comprise the cytosolic domain (white), transmembrane domain (wavy lines), signal sequence (dark grey), GH9 catalytic domain (light grey), linker region (thick black line), and carbohydrate binding module (dots). Structural subclasses are represented by SlGHl9A1/TomCel3 (class A, U78526), SlGH9B1/TomCel1 (class B, U13054), and SlGH9C1/TomCel8 (class C, AF098292). To standardize the nomenclature for these gene families, we suggest the following, using genes encoding GH9 enzymes as an example: an indication of the genus and species, followed by the designated the GH family (GH9). The letters A to C adjacent to the family number correspond to the domain structure, or subclass, of the corresponding protein (Fig. 1), which will provide additional information about potential function. We note that this nomenclature helps determine the structural subclass with which a particular GH9 gene is associated, rather than specifically suggesting an orthologous sequence. As an example of the naming scheme, we have applied the guidelines to rename the members of the GH9 family from tomato (Solanum lycopersicum), the plant species from which the greatest number of family members has been studied in detail. Historically, members of the tomato GH9 family have been referred to as TomCel1-8 and their new designations are shown in Table I GH9 proteins of tomato Features noted are cytosolic domain (CT), transmembrane domain (TM), signal peptide (SP), glycosyl hydrolase family 9 catalytic domain (GH9), and a family 49 carbohydrate binding module (CBM49). Peptide fragment. GH9 proteins of tomato Features noted are cytosolic domain (CT), transmembrane domain (TM), signal peptide (SP), glycosyl hydrolase family 9 catalytic domain (GH9), and a family 49 carbohydrate binding module (CBM49). Peptide fragment. GH9 proteins of Arabidopsis Features noted are cytosolic domain (CT), transmembrane domain (TM), signal peptide (SP), glycosyl hydrolase family 9 catalytic domain (GH9), and a family 49 carbohydrate binding module (CBM49). Nicol et al. (1998). Mølhøj et al. (2001). Shani et al. (1997). Yung et al. (1999). del Campillo et al. (2004). GH9 proteins of Arabidopsis Features noted are cytosolic domain (CT), transmembrane domain (TM), signal peptide (SP), glycosyl hydrolase family 9 catalytic domain (GH9), and a family 49 carbohydrate binding module (CBM49). Nicol et al. (1998). Mølhøj et al. (2001). Shani et al. (1997). Yung et al. (1999). del Campillo et al. (2004). This nomenclature conforms to that used for bacterial endoglucanases with some slight modifications, since many of the families of plant GHs are much larger than those of bacteria (CAZY) and so the information designating the order in which they are reported is more easily presented by a numerical rather than an alphabetical system. For example, in Arabidopsis there are 49 members of family GH1, 33 of GH16, and 69 of GH28, all of which contain more members than there are letters in the alphabet (http://www.cazy.org/geno/3702.html). The new nomenclature has been presented in the context of the plant GH9 family, but we further suggest that similar guidelines be adopted for naming members of other GH families and that future efforts to standardize nomenclature be coordinated in consultation with CAZY. Similarly, the use of the established naming schemes for carbohydrate binding modules (www.cazy.org; Boraston et al., 2004) will further help advance the study of analogous hydrolytic enzymes and other associated functional domains from different organisms, both within and between taxa.
No takes yet. Share an insight, caveat, or question.
Urbanowicz et al. (2007) studied this question.