Relationships between supracellular components (biological systems), intracellular components, and the function and behavior of these components are revealed by the interaction of individual components. Systems biological approaches aim at modeling these interactions to find primary relationships and to distinguish causality and effect. The understanding of how these interactions are regulated allows making predictions on function, behavior, and survival. Transcript profiling offers the largest coverage and a wide dynamic range of gene expression information and can often be performed genome wide. Microarrays are currently most popular for transcript profiling and can be readily afforded by many laboratories. Various commercial and academic microarray platforms exist that vary in genome coverage, availability, specificity, and sensitivity (Table I Advantages and disadvantages of various technologies for the measurement of transcript and protein abundance A systematic performance assessment for the different protein quantification techniques was recently conducted (Turck et al., 2007) and a detailed description of the different quantification techniques along with examples for application in the plant field is available (Baginsky, 2009). Advantages and disadvantages of various technologies for the measurement of transcript and protein abundance A systematic performance assessment for the different protein quantification techniques was recently conducted (Turck et al., 2007) and a detailed description of the different quantification techniques along with examples for application in the plant field is available (Baginsky, 2009). Much control of gene expression occurs at the level of transcription, and information on genome-wide chromatin profiles (epigenomes) and transcription factor binding to promoters is needed to decipher the inherent logic of transcriptional regulation. Chromatin immunoprecipitation (ChIP) coupled to microarray analysis (ChIP-chip) or high-throughput sequencing (ChIP-Seq) can generate such data. In plants, DNA methylation, repressive and activating chromatin marks, as well as histone variants have been mapped onto the genome (for review, see Zhang, 2008), but because such marks are expected to differ between cell types and developmental stages, more targeted epigenome profiling is needed in the future. Targeted analysis of DNA methylation during seed development, for instance, revealed unexpected genome-wide demethylation (Gehring et al., 2009; Hsieh et al., 2009). ChIP-chip was also used for global mapping of binding sites of transcription factors such as TGA2 and SEPALLATA3 and to refine definitions of binding motifs that were previously determined by in vitro experiments (Thibaud-Nissen et al., 2006; Kaufmann et al., 2009). It was found that SEPALLATA3 is a key component in the regulatory transcriptional network underlying the formation of floral organs. In a comparative experiment ChIP-chip and ChIP-Seq gave very similar results (Kaufmann et al., 2009). This is encouraging because bias introduced by the profiling technology seems not to severely confound studies on global protein-binding profiles. Currently, work is going on in several laboratories to establish a compendium of transcription factor binding sites in Arabidopsis (Arabidopsis thaliana). Thus, more genome-wide data sets are in reach that could provide causal explanations for transcriptional profiles. Gene expression is a highly regulated, multistep process, and it is impossible to predict the exact protein concentration or activity from the measurement of mRNA levels. Proteomics has therefore become a key tool in systems biology because it provides quantitative and structural information about proteins, which are the major functional determinants of cells. Phenotypic alterations associated with genetic perturbations often result from changes in protein accumulation or stability, or changes in protein posttranslational modifications, which can disrupt protein-protein interactions and network connectivity (Gstaiger and Aebersold, 2009). Quantitative protein information complements data from transcriptional profiling and metabolomics. It represents a key link between different levels of gene expression regulation and provides insights into their causal relationships. Unlike transcriptional profiling, however, comprehensive proteome analysis remains challenging, and information about proteome complexity and dynamics is far from complete (Cox and Mann, 2007). Moreover, the rate of metabolite synthesis is often controlled by regulatory posttranslational modifications of enzymes and not only by their abundance. Information about quantitative relationships between RNA and protein accumulation, posttranslational protein modifications, and metabolite levels is therefore required to fully understand regulatory circuits that control systems behavior and function. Protein quantification can be absolute or relative (Table I). While relative protein quantification mostly depends on stable isotopes, absolute quantification of comprehensive protein sets is much more difficult. Recent improvements in statistical data evaluation and increasing accuracy of mass spectrometry instruments allow quantifying large numbers of proteins in shotgun-type experiments on the basis of spectral counting (Lu et al., 2007). This method is reliable and comparable to most other quantification methods, including two-dimensional PAGE-based protein staining; however, the protein dataset must be very large. More accurate information about the exact in vivo concentration of individual proteins requires specialized targeted approaches. Current methods for absolute protein quantification include isotope dilution strategies using isotopically labeled peptides as internal standards (for a comprehensive review, see Brun et al., 2009). Signature peptides for internal standardization are characteristic for a protein of interest, and are often referred to as proteotypic peptides (PTPs). In AQUA, PTPs are added to analytical protein samples in known concentrations. The protein samples are subsequently scanned for PTPs of interest. Using the extracted ion chromatograms the native peptide can then be quantified relative to the added PTP (Kuster et al., 2005). A modification of this strategy accounts for quantification errors derived from incomplete tryptic digest of the analytical sample. In QconCAT (for quantification concatamer), a synthetic protein with concatenated, isotopically labeled PTPs is expressed as recombinant protein in a biological system, added to the sample prior to Trypsin treatment and carried through the digestion procedure, such that losses from incomplete tryptic digestion will also affect the quantity of the PTPs. Both the AQUA and the QconCAT strategies are incompatible with upstream fractionation techniques, which is a potential problem in biomarker quantification. A way around this constraint is offered by the protein standard absolute quantification strategy, which uses isotopically labeled protein standards that are added to the sample prior to fractionation. Several prediction tools exist that help to define the most suitable PTPs for the detection and quantification of specific proteins. However, only experimental data provide the necessary reliability for PTP selection because in practice PTP prediction often deviates from experimental observations. Therefore, efforts are under way to catalogue PTPs for model organism proteomes. Proteome maps for Arabidopsis generated PTPs for 4,105 proteins, many of which may be optimal for the detection of proteins in different organs (Baerenfaller et al., 2008). Similar quantitative approaches are also used for metabolites, because in addition to RNA and protein levels, understanding the function and behavior of metabolic networks requires global information about metabolite concentrations and fluxes as well. In recent years, much progress has been made in metabolic profiling, and the interested reader is referred to recent reviews (e.g. Issaq et al., 2009, and refs. therein). During the analysis of large gene expression datasets the researcher is often confronted with several questions. How do we interpret a mathematical relationship between genes or between genes and conditions? For example, does a high correlation between two genes mean that they are coregulated, or could one of them be the positive regulator of the other? Or can we assume that they are involved in the same pathway or biological process? Although it is not possible to answer these questions conclusively from gene expression data alone, a number of parallel approaches can be useful to distinguish between different scenarios. For example, Gene Ontology enrichment analysis can provide confidence that a given gene cluster is enriched in genes that are known to have a common function, cellular location, or biological process. Similarly, conserved cis-regulatory elements in the promoters of genes from the same cluster indicate that they are likely coregulated. Although these methods do not establish proof of the nature of the relationship between genes, they allow formulating hypotheses that can be tested in the laboratory. In summary, although gene expression analysis by itself is rather descriptive (i.e. describing how genes respond to various test conditions or tissues), it is a valuable validation tool and an excellent starting point to study novel cellular process and to formulate novel hypotheses. A major challenge of genome-scale transcription analysis is the very large number of predictors (genes) compared to a generally small number of measurements (microarrays). Without appropriate statistical measures to correct for multiple testing and including false discovery rates, almost any approach will yield significant genes, including many false positives. The creation of large databases in recent years has brought an additional layer of complexity and precautions to take (see Table II Overview of some of the most popular plant gene expression microarray platforms and the number of available experiments in ArrayExpress The Arabidopsis ATH1 array is the most frequently used microarray, followed by the CATMA 25k and 23k arrays. In all, approximately 750 Arabidopsis microarray experiments have been published so far. Rice (Oryza sativa) and barley (Hordeum vulgare) are the second and third plant species in terms of microarray experiments published. Soybean (Glycine max) also has a high number of arrays, but this is due to a single very large experiment containing 2,521 arrays. IPK, Leibniz Institute of Plant Genetics and Crop Plant Research; TIGR, The Institute for Genomic Research. Overview of some of the most popular plant gene expression microarray platforms and the number of available experiments in ArrayExpress The Arabidopsis ATH1 array is the most frequently used microarray, followed by the CATMA 25k and 23k arrays. In all, approximately 750 Arabidopsis microarray experiments have been published so far. Rice (Oryza sativa) and barley (Hordeum vulgare) are the second and third plant species in terms of microarray experiments published. Soybean (Glycine max) also has a high number of arrays, but this is due to a single very large experiment containing 2,521 arrays. IPK, Leibniz Institute of Plant Genetics and Crop Plant Research; TIGR, The Institute for Genomic Research. Most transcript and protein profiling experiments analyze mixtures of tissues containing different cell types and organelles. This approach reveals certain global patterns, but quantitative analyses and modeling is limited with such complex data. Therefore methods for organ (or better) cell-type-specific transcript and protein profiling as well as for organelle-specific proteomics are needed. Four types of approaches are now commonly used to sample RNA and/or proteins from selected cell types: (1) micropipetting, (2) laser capture microdissection (LCM), (3) protoplasting and sorting, and (4) polysome immunopurification (for review, see Zanetti et al., 2005; Hennig, 2007; Nelson et al., 2008). Micropipetting using microcapillaries directly extracts the contents from selected cells. It has been successfully applied to various leaf cell types and for phloem but extraction is more difficult from internal cells. LCM involves sectioning of frozen or embedded tissue, and subsequent dissection of the region of interest using laser excision. Applications of LCM include studies of vascular tissue, epidermis, and pericycle in maize (Zea mays) and seed development in Arabidopsis. Micropipetting and LCM are usually very labor intensive and difficult for isolation of small cells such as in meristems. Because of the limited amount of material that can be captured, they work well for transcript profiling, which can use amplification steps, but provide only a very small coverage of the proteome. As an alternative, protoplasting and cell sorting offers rapid and accurate isolation of RNA from small cells. Specific tissues or cell types that are labeled by expression of GFP are isolated by protoplasting and sorted through a fluorescence-activated cell sorter. Millions of cells can be processed within 1 to 2 h, but care has to be taken to exclude changes in gene expression profiles by sample processing. This technique was successfully applied to measure genome-wide expression profiles in more than 15 root regions, establishing a compendium of digital in situ data (Birnbaum et al., 2003; Cartwright et al., 2009). It will be interesting to test whether this approach can also be used for protein profiling. Polysome immunopurification is based on the tissue-specific expression of the FLAG-tagged ribosomal protein L18 in transgenic plants (Zanetti et al., 2005). In contrast to micropipetting, LCM, and sorting of protoplasts, which all can be used to isolate total cellular RNA, polysome immunopurification can be used to isolate transcripts that are associated with ribosomes (translatome). Discrepancies between total RNA levels and representation translatome can reveal regulation at the level of translation (Mustroph et al., 2009). In the future, translatome datasets, which bridge transcriptomics and proteomics, can help to interpret unusual transcript-to-protein ratios (see below). Alternatively, it is possible to identify cell-type-specific transcripts and proteins by comparing wild-type plants with mutants that lack specific cells or tissue types. In Arabidopsis, for instance, a series of homeotic mutants that lack various floral organs was used to identify several hundreds of floral organ-specific genes (Wellmer et al., 2004). If no appropriate mutants exist, specific cell types can be genetically ablated by expression of a cell-autonomous toxin, such as diphtheria toxin subunit A or RNase, under the control of cell-type-specific promoters. Again, these approaches have been proven to work for transcript profiling (Tung et al., 2005) but it remains to be tested whether they could be useful for protein profiling. Systematic analysis of accurate protein localization is essential to understand cellular networks in the context of compartmentalization, which is a fundamental design principle of eukaryotic cells. Organelle proteomics has therefore become a very active research field. Until recently, the protein inventory of cell organelles was based on proteins from isolated organelles, such as mitochondria, chloroplast, and peroxisomes (Lilley and Dupree, 2007; Baginsky, 2009). This approach has limitations because true low-abundant organelle proteins often cannot be distinguished from contaminating proteins. Two approaches have been used to deal with this problem. First, a recently reported isolation procedure for mitochondria used the electrostatic characteristics of the mitochondrial surface to separate mitochondria from other organelles in an electric field. This procedure results in mitochondria preparations with higher purity, but the yield is low (Eubel et al., 2007). Second, information about the quantitative distribution of proteins along density gradients has been used to determine if a protein was enriched by the organellar isolation procedure. In practice, the abundance distribution profile of unknown proteins is compared to known organelle marker proteins. This strategy is referred to as protein correlation profiling (Foster et al., 2006) or LOPIT (Dunkley et al., 2006). Both procedures, however, are of limited use for the analysis of proteome dynamics in response to a stimulus because the long time that is needed to isolate and purify organelles affects their proteome properties. This is especially critical for transient posttranslational protein modifications. Thus, proteome dynamics is best analyzed at the cell or tissue level, followed by sorting of proteins into their respective organelle a posteriori. This strategy is now possible because substantial information about the protein complement of different cell organelles has accumulated (a comprehensive collection of proteome databases is for example available in Lu and Last, 2009). The SUBA database is most suitable for this purpose, because it is frequently updated and well maintained. SUBA generates lists of organelle proteins using reliability criteria, for example evidence from several different proteomics studies, targeting prediction, or GFP-localization assays, or a combination of this information (Heazlewood et al., 2007). For the chloroplast, two proteome reference tables have been established (Yu et al., 2008; Reiland et al., 2009). The overlap between these two proteome reference tables has generated a list of 1,156 proteins that can be considered high-confidence chloroplast proteins. Although the number of organelle proteins is constantly increasing, it is not clear when an organelle proteome can be considered complete. Organelle proteomes are dynamic and functional organelle proteomes differ significantly during development, in different cell types or tissues, and in different conditions. This problem can be addressed by considering organelles as cellular subnetworks and applying flux-balance modeling to assess network consistency. Initial modeling approaches with mitochondria and chloroplasts focused on a limited number of reactions, such as those of the Calvin cycle, amino acid biosynthesis, or the tricarboxylic acid cycle. Also, mitochondrial network reconstructions based on proteomics data are available and the existing models allow prediction of metabolite accumulation for a limited number of metabolites (Vo and Palsson, 2007). A recent flux-balance model of the primary metabolism in into mitochondria, and the and the of different organelles to and 2009). The examples the excellent of metabolic network to identify in existing analysis of transcript and protein abundance in Arabidopsis based on genes from various primary and metabolism Transcript abundance was as a expression derived from multiple ATH1 array measurements from leaf samples from et al., 2008). The proteome data was from leaf of these ratios of protein to transcript abundance from discovery is a describing the of the nature of relationships between and associated from of a biological types of networks have been with to the types of involved and the of the network or small While analyses aim at novel of the global networks usually additional data types that cannot be on a global and use models that allow a more prediction of network A more recent development is the of various networks into an of in which networks are considered as strategies that and for et al., 2007). network reconstructions have increasing in the years and several genome-scale models are now available for and However, of the existing models can currently provide a of all in a cell For example, the most genome-scale models for only approximately of the approximately metabolic genes and only approximately of the are functional et al., 2009). The is for higher and for plants no genome-scale model to although efforts to databases are under way et al., 2008). in metabolic network depends on the functional of the genes that proteins with unknown which is the for about of the Arabidopsis plant metabolic network mostly on specific such as acid synthesis or metabolism et al., 2009; et al., 2009). Systematic efforts to genome are under and the is well The Arabidopsis Information which and the of material and example of such a functional is the that was at (Lu et al., 2008). for all known chloroplast proteins are generated and also at the metabolite In a in the functional of genes is key to networks with and consistency. Similar efforts are under way to plant transcriptional regulatory for example those that control and root development et al., et al., 2007; 2008), or the et al., 2006). this in are used in of genetic regulatory such as are for a small number of genes and have been used to model the pathway network from data to genes associated with this network et al., 2004). A similar modeling was carried on a of genes in Arabidopsis and to novel components in various networks et al., 2007). In the future, analysis of transcriptional regulatory networks to also epigenome and transcription factor binding data from ChIP-chip and ChIP-Seq network is also used in plant For example, et that for many genes in expression could be by expression quantitative using recombinant of Arabidopsis. expression quantitative mapping and regulator gene gene regulatory networks for time could be that were in with published data. The combination of data with quantitative data is expected to the understanding of complex regulatory networks such as and and A of models exist for mapping complex and to changes in et al., 2006). to all work could not be due to
No takes yet. Share an insight, caveat, or question.
Baginsky et al. (2009) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: