Key points are not available for this paper at this time.
New ways of capturing and representing biological knowledge are needed to enable individual researchers to remain abreast of relevant discoveries and to permit computational approaches for interpreting the large volumes of diverse data generated by modern biological research. Here, we describe a promising approach that expands the term “reaction” to represent biological processes. We show how users can represent a wide variety of biological processes in plants in terms of the concept of a reaction and assemble the information obtained from the model plant Arabidopsis thaliana into an online knowledgebase called Arabidopsis Reactome. Its curated and imported pathways currently cover ∼8% of the Arabidopsis proteome. Arabidopsis Reactome events have also been electronically projected onto five other predicted plant proteomes. Such a system allows the visualization and interpretation of high-throughput data, hypothesis formulation in systems biology, and is a useful learning resource. The Arabidopsis Reactome project (www.arabidopsisreactome.org) is open access, open source, and open to contributions. Currently, the genome sequences of six higher plants and a moss species have been assembled, annotated, and published (Arabidopsis Genome Initiative, 2000; International Rice Genome Sequencing Project, 2005; Tuskan et al., 2006; Jaillon et al., 2007; Velasco et al., 2007; Ming et al., 2008; Rensing et al., 2008). The availability of large populations of sequence-tagged insertion mutations for nearly all Arabidopsis genes (Alonso and Ecker, 2006), surveys of polymorphisms in many Arabidopsis ecotypes (Clark et al., 2007), and the free availability of microarray data and data-mining tools (Craigon et al., 2004; Zimmermann et al., 2004) have all greatly accelerated the scale and scope of plant research (Somerville and Koornneef, 2002), reflected by more than 240 publications per month citing Arabidopsis. Given the acceleration in the number of plant genome sequencing projects, it is increasingly difficult for individual researchers to stay abreast of relevant literature and to make connections between different sets of information. This difficulty is compounded by the general inaccessibility of information contained in the literature to computer-based analysis, which severely limits its value as a source of biological knowledge (Jensen et al., 2006). Therefore, developing computational methods for capturing and representing biological knowledge is a high priority, particularly for model organisms that are the focus of most experimental work. Several bioinformatics resources and software packages have been developed to manage and exploit the wealth of data generated by plant genome projects, functional genomics resources, and high-throughput transcriptomics experiments. The predicted proteins in the Arabidopsis genome have been systematically described by The Arabidopsis Information Resource (TAIR; http://www.arabidopsis.org/portals/genAnnotation/) using Gene Ontology (GO)–controlled descriptions of gene functions according to the biological process, molecular function, and cellular component of individual genes (Ashburner et al., 2000). This has enabled much more rapid, accurate, and consistent assignment of predicted functions to genes and permits the development of more accurate relationships between genes in different organisms. Several databases and software applications relate gene entities to each other in networks in Arabidopsis. The AraCyc database (Mueller et al., 2003) displays computationally predicted Arabidopsis metabolic pathways that are largely manually curated. MAPMAN uses a hierarchical ontology different from GO terms that can be used for visualizing large data sets onto metabolic pathways and other biological processes (Thimm et al., 2004). The VirtualPlant (Gutierrez et al., 2007) and ONDEX (Kohler et al., 2006) systems have created graph-based integrations of knowledge and gene functional inferences that may be queried, filtered, and appended using tools like Cytoscape (Suderman and Hallett, 2007) to generate new functional insights. GENEVESTIGATOR (Zimmermann et al., 2004, 2005) provides web-based analytical services that relate gene expression data to a wide variety of gene-related entities, such as GO terms, mutant phenotypes, pathways, and developmental processes. Reactome is an extensively curated pathway knowledgebase that focuses on human processes (Joshi-Tope et al., 2003, 2005; de Bono et al., 2007; Vastrik et al., 2007). A key feature of Reactome is its elegant data model that extends the notion of a biochemical reaction, where substrates go in, products come out, and a catalyst is frequently required to lower the free energy of the transformation. This concept also can be used to represent the binding of a ligand to a membrane receptor, the formation of a complex, the binding of a transcription factor in a promoter region, or the translocation of a molecule between subcellular compartments. In this way, the data model expresses molecular processes in the same way that scientists understand them and allows connected reactions to represent biological processes (e.g., transcription and cell cycle) in terms of their underlying molecular transformations, associations, and translocations. Based on this extended definition of reaction, reactants and products can be proteins, lipids, nucleotides, small molecules, or complexes of these. The data model further distinguishes among the topologically or functionally different forms of each molecule. For instance, this allows the distinction between chloroplastic maltose and cytosolic maltose or between the various posttranslational modifications of a protein. Here, we describe Arabidopsis Reactome, a knowledgebase of biological processes from the model plant Arabidopsis. Release 2 (www.arabidopsisreactome.org) comprises seven curated and 311 imported superpathways that together represent 8% of the Arabidopsis proteome. We show, using examples based on the mitotic cell cycle, that the knowledgebase has wide applicability for exchanging structured data with other databases, for comparative network analysis, data integration, and for visualization and protein interaction analysis. The straightforward authoring tool and realistic detailed descriptions of biological processes inherent in Reactome's data model provide an excellent foundation for representing, exchanging, and integrating biological information, suggesting that it will find wide application in the Arabidopsis community as a gold standard for pathway knowledge and key foundation for systems biology research. The information in Arabidopsis Reactome was generated from curated pathways that have been manually entered and reviewed by experts and imported pathways from several third-party pathway databases. Knowledge acquisition for the curated pathways followed the process established for the human Reactome system. Essentially, pathways were authored, curated, and peer reviewed by expert biologists (PhD level and above) and bioinformaticians. The curatorial process used a set of applications, namely, Reactome Author Tool and Reactome Curator Tool, developed specifically for the purpose of collecting and validating pathway models (Joshi-Tope et al., 2005). Every protein, gene, or small molecule in Arabidopsis Reactome has a reference identifier that points into a public reference database. In the case of protein sequences, the primary source of identifiers is UniProt. Entities such as chemical compounds are referenced by the ChEBI database. For the imported KEGG and AraCyc enzymes and chemical compounds, referencing was performed in an automated manner that is also applicable to newer additions. In addition, entries were automatically cross-referenced to external databases, such as UniProt (Schneider et al., 2005), TAIR (Rhee et al., 2003), Munich Information Center for Protein Sequences (MIPS) (Schoof et al., 2004), National Center for Biotechnology Information (NCBI) Entrez Gene (Maglott et al., 2005), KEGG COMPOUND (Kanehisa and Goto, 2000), and ChEBI (Degtyarenko et al., 2007). The long-term goal in Arabidopsis Reactome is to establish a detailed set of curated pathways representing all major biological processes in Arabidopsis. Initially, metabolic pathways and the mitotic cell cycle were selected as contrasting processes that would challenge the plasticity of the data model and provide a foundation for data integration and modeling. An essential piece of information for a reaction to be incorporated into the curated area of Arabidopsis Reactome was the existence of experimental evidence, usually a reference to a published article. In addition to reactions and literature references, the data model contains fields for species, GO molecular function, subcellular location, and other relevant information that were filled out during the curatorial process. In some instances, Arabidopsis reactions imported from KEGG and AraCyc databases were used as structured reference material to start the curatorial process. Overview of the Arabidopsis Reactome Home Page. The panel with the arrows is the reaction map. The arrows represent reactions that have been manually created based on evidence from the literature and reactions that have been imported from external databases (AraCyc and KEGG). Below the reaction map is the table of modules that contains a list of all the superpathways that are present in Arabidopsis Reactome. We imported Arabidopsis metabolic pathways from KEGG (release 38.0) and AraCyc (release 2.5) databases into Arabidopsis Reactome as text files from their ftp servers. The files were parsed, and the data were stored in a MySQL relational database using custom database schemata developed to represent each source. Using the Perl XML∷Generator module (http://www.cpan.org/), these data were used to construct documents in the Reactome Author Tool native, XML-based, GKB format. The documents were then opened with the Author Tool, and reactions were joined manually to form pathways according to the pathway diagrams found at their source websites. These files were then imported into the Reactome Curation Tool that was used to deposit the data in the central Arabidopsis Reactome database. Once in the database, the arrows representing the reactions were manually laid out in the reaction map using the Reactome pathway visualization tool. Pathways involved in related or similar processes were laid out in close proximity to each other within either the KEGG or AraCyc areas of the reaction map. Importation software has been written to allow Arabidopsis Reactome to be updated from KEGG and AraCyc sources as new versions become available. View of a Selected Reaction in Arabidopsis Reactome. The image shows a diagrammatic representation of the reaction components and a list of information about the reaction, such as the inputs, outputs, catalyst, GO molecular function, preceding and following events, compartment, literature references, and equivalent events in other organisms. A hierarchical structure of the involved pathway can also be viewed by expanding the “event hierarchy” frame (shown on the left). SUBA (Heazlewood et al., 2007) and UniProt subcellular localization information was used to assign the subcellular location of the imported reactions. For example, the enzyme peroxisomal 2,4-dienoyl-CoA reductase was used to assign its catalyzing reaction “monovinyl protochlorophyllide a + NADPH NADP+ + chlorophyllide a AT3G12800” and its components to the peroxisome. This substantially enriches the knowledge associated with metabolic pathway data. The predicted proteomes from the published genomes of rice (Oryza sativa; International Rice Genome Sequencing Project, 2005), poplar (Populas trichocarpa; Tuskan et al., 2006), the moss Physcomitrella patens (Rensing et al., 2008), and the two grape varieties (Vitis vinifera and V. vinifera var Pinot Noir; Jaillon et al., 2007; Velasco et al., 2007) were downloaded from their websites. Using the NCBI BLASTP algorithm, we matched the Arabidopsis proteome downloaded from the TAIR ftp site to the predicted proteomes of these species and then fed the results to the OrthoMCL algorithm (Li et al., 2003) to identify and cluster the orthologous proteins into groups. OrthoMCL results were then appropriately formatted so they could be used by the scripts included in the Reactome system to produce the equivalent organism-specific reactions and pathways. A total of 8269 reactions and 2196 proteins from Arabidopsis were projected onto the five plant species (Table 1 Statistics from the Orthologous Transfer of Arabidopsis Reactions onto Five Other Sequenced Plant Genomes Reactions and pathway figures include duplicated events found in more than one superpathway. Statistics from the Orthologous Transfer of Arabidopsis Reactions onto Five Other Sequenced Plant Genomes Reactions and pathway figures include duplicated events found in more than one superpathway. Comparison of the Arabidopsis cell cycle with the electronically inferred cell cycle in rice, poplar, P. patens, and the two grape varieties showed that almost 60% of the reactions were conserved between Arabidopsis and poplar, compared with 45% in rice and 33% in moss (Table 2 Projection of the Arabidopsis Mitotic Cell Cycle Reactions onto Five Sequenced Plant Genomes Projection of the Arabidopsis Mitotic Cell Cycle Reactions onto Five Sequenced Plant Genomes The Arabidopsis Reactome home page is divided into two main panels: the reaction map and the table of contents (Figure 1). The reaction map displays all the reactions contained in Arabidopsis Reactome in the form of arrows. Arrows joined together represent biological pathways, and pathways are clustered according to related biological processes. The left side of the reaction map is occupied by Arabidopsis reactions found in KEGG (Kanehisa et al., 2002), and the right side is taken by those reactions found in AraCyc (Mueller et al., 2003), leaving the center for those reactions that have been manually curated. The table of contents is seen under the reaction map. It lists all the primary pathways (superpathways) present in Arabidopsis Reactome, starting with the curated pathways at the top of the table and followed by pathways from KEGG and AraCyc. As curated, peer-reviewed pathways are added to the central part of the reaction map, equivalent pathways from KEGG and AraCyc are removed, increasing the number of reactions in the central panel and decreasing those from the side panels. Entries in Arabidopsis Reactome can be searched using simple or advanced search facilities located at the top of the home page. By selecting a reaction, the web interface of Arabidopsis Reactome returns an increasing level of detailed information. This includes a description, the reaction components (compounds, enzymes, etc.), GO annotation (subcellular location and molecular function), preceding and following events, organism name, equivalent events in other organisms, reference in the literature, and, where applicable, links to external databases, such as UniProt, TAIR, MIPS, NCBI Entrez Gene, KEGG COMPOUND, and ChEBI (Figure 2). A key feature of a curated knowledgebase is its utility for integrating diverse data sets into a comprehensive description of biological processes. SkyPainter is a useful feature of the Reactome system that allows researchers to visualize and analyze their own data sets in relation to the reaction maps. It can be found on the top menu bar of the Arabidopsis Reactome home page. Researchers can upload a list of genes or other identifiers to color the reaction map in a number of The SkyPainter module a of gene such as Arabidopsis Genome and UniProt with electronically inferred pathways can also be searched by selecting the plant species on the SkyPainter page and using the gene or We the of SkyPainter using two sets of published experimental for the of genes (Li et al., 2006) and for the visualization of the of gene expression during the cell cycle et al., SkyPainter of A total of genes were by at compared with at SkyPainter genes within The is in two followed by an list of pathways the top is the reaction map the reactions according to the number of genes in a is a list of pathways according to the the of the number of genes in a pathway by A shows the genes to be in the SkyPainter of a Cell Cycle Gene expression was at points in Arabidopsis for cell cycle SkyPainter the reaction arrows according to the of the expression of all components in the at the shows the reactions within a of gene expression can be viewed in 1 and Protein with the Cell Cycle The Arabidopsis Reactome cell cycle (shown as the of arrows to represent the different of the cell cycle) was with the experimental cell cycle protein (shown in the center of the cell from were also added to the network as within the The algorithm was then used to of proteins that could with the cell The top seven are and according to the of the cell cycle they The are as and and binding and is in in and the experimental cell cycle a in all of the cell The are proteins of the proteins have more than one in the Arabidopsis Reactome cell cycle, the protein is part of a a is to one reaction for The data contained in Arabidopsis Reactome can be viewed in Cytoscape et al., 2003) and or can be in level 2 and 2003), and format. The of the Arabidopsis Reactome and tools for biological pathways can be downloaded from the Arabidopsis Reactome by following the has been a in the availability of pathway tools the focus of biological research has a systems of biological processes. The of a pathway tool are of of data model and experimental evidence, wide standard facilities to data and that it is open source and open AraCyc and KEGG are comprehensive and pathway resources from which we have incorporated the data into Arabidopsis Reactome. AraCyc has the of metabolic pathways, many of which have been manually curated. The standard that these pathways are used in other pathway resources, such as et al., 2007) and Plant (Gutierrez et al., 2007). KEGG is by the data and 2005) and are sources of curated pathway information that can be used to the of pathways in the the of that these pathways be automatically incorporated into Arabidopsis Reactome. MAPMAN (Thimm et al., 2004) is a tool for large genomics data sets onto diagrams of metabolic pathways and other processes. Its wide of metabolic and other biological expert and useful hierarchical with gene make it a for integrating it a standard for exchanging pathway information with other pathway resources, and it the level of of reactions compared with Reactome to its Currently, we are developing a between Arabidopsis Reactome and The human Reactome has established a data model and software tools that allow scientists to a wide variety of biological knowledge in a form computational and these have been developed as a knowledgebase for human biological processes (Joshi-Tope et al., 2005). from the human Reactome, are currently other under development for and In this we have the application of the Reactome system to the reference plant species Arabidopsis and how the data model many of different biological processes. these models currently a small of the molecular entities in this is an a network The of Arabidopsis Reactome the functions of proteins curated from primary literature These proteins in a number of pathways, the mitotic cell cycle, cell cycle cell and several metabolic pathways. proteins have been imported from KEGG (Kanehisa and Goto, and AraCyc (Mueller et al., 2003) pathway databases. with the curated proteins, they represent 8% of the Arabidopsis genes The curated and imported proteins in Arabidopsis Reactome in a total of 8269 reactions from The Arabidopsis proteome was electronically projected in Arabidopsis Reactome onto five other published plant and moss predicted proteomes orthologous This allows the visualization of the equivalent reactions and pathways on other plant species and pathway between Researchers can with the web-based interface to pathways and visualize data on the reaction map. Arabidopsis Reactome pathways allow users to based on a much of data and knowledge than These include of network et al., 2007), network such as 2007) the of system 2005; and 2008), metabolic and 2002), and the to used for the of networks from data et al., 2007). The addition of functions and their will enable approaches to and metabolic and We have examples of this orthologous from Arabidopsis to the of genes that are involved in the same pathway in other species and in gene between plant Such can be using functional genomics approaches and gene expression data on pathways using SkyPainter genes within pathways and gene expression data in a and by curated Arabidopsis Reactome such as the cell cycle with experimental et al., 2007) and predicted et al., 2007), we of genes that can be for cell It is that other (e.g., interaction or protein could also be used as a or in the Cytoscape and with a network such as for example, to establish interaction of the knowledge of biological processes in plants is described as free text in the published literature and other The Reactome approach is to this knowledge using and understand the and can make expert on evidence and can in the of This process a high of knowledge and for reference species such as it will the plant research community such pathway knowledge can be projected electronically onto other plant genomes using Reactome. Currently, knowledge from Arabidopsis has been electronically to poplar, rice, and P. patens and is from the Arabidopsis Reactome it be to plant genomes by orthologous from Arabidopsis Reactome using such as OrthoMCL (Li et al., 2003) and et al., As plant are currently this of Arabidopsis Reactome will be methods of knowledge the of text and that are largely to is currently of this knowledge in a computational form (Jensen et al., it is currently an to the process than a comprehensive way to the and of Arabidopsis Reactome among users is a model the process and Such an approach has been that will to deposit data in the TAIR database as part of the process Arabidopsis Reactome is for this of as it a wide variety of knowledge about biological entities, such as proteins, proteins, and This and other examples would its and the of electronically knowledge would be more The of pathway knowledge representation is in pathway as it the integration of data in systems biology to generate new et al., 2007). Arabidopsis Reactome provides a based on and comprehensive of the literature, which will many areas of biological research in plants and to biological The following are in the online of this article. SkyPainter of Arabidopsis Reactome Cell Cycle with the Cell Cycle in Cell Cycle with of Protein 1 to Using and in the Cell Cycle of within Using the Gene from the SkyPainter Gene during the Cell We the Reactome development for their and useful and for the microarray data and analysis. This was by from the Biotechnology and under the and and by the Arabidopsis integrating and are by a from the National of a from the and from the National of Cell and the
Tsesmetzis et al. (2008) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: