Since its inception in 1989/1990, the Human Genome Project has put major emphasis on structural genome analyses, resulting in elaborate genetic and physical maps. As a result of these achievements, important disease genes have been identified through positional cloning, and the focus of the project is now turning toward sequencing the entire genome. Complete sequences of genome DNA have already been determined for several single cell organisms [1-3]including Saccharomyces cerevisiae (yeast) [4], and vigorous analyses are being carried out on many other organisms [5, 6]. Sequencing the human genome is expected to be complete around the year 2005. The goal of genome analysis is to decode the entire genetic information carried in the genome. This means that, besides sequencing, information which cannot readily be read from structural data must be collected. Examples are the activities of hypothetical gene products, the expression control and the regulatory network of each gene. This requires systematic efforts which may be collectively called functional analysis of the genome. In this connection, a consortium has been established for functional analysis of the S. cerevisiae genome through systematic disruption of genes [7] For higher eukaryotes with more genes, however, different strategies are needed, such as enlisting all (or nearly all) the active genes and studying the sites of their expression together with the extent of their activities. A list of active genes and their activities in the cells and tissues can be called an expression profile. Construction of such a list should be started with the representative cells or tissues of the body and expanded to cover the various stages of development or pathological conditions. Kohara et al. have initiated the collection of large numbers of in situ staining patterns with Caenorhabditis elegans, using probe cDNAs that represent novel genes as discovered by single run cDNA sequencing. This worm is transparent and is composed of only 965 cells, thus allowing for whole body analysis [8, 9]. The facts that cell lineages are well established during development and that there is only a small number of genes for testing favor adoption of this procedure. In the case of man, body size and the large number of genes preclude a similar approach. Thus, Okubo et al. have initiated efforts to identify active genes by single pass sequencing of cDNA obtained from a type of cell or tissue and quantified their activities in the mRNA population [10]. Although an adult human body consists of some 6 trillion cells, they can be categorized into 200 basic types [11]. Therefore, the number of cells or tissue species is within a reasonable range. The cDNA libraries used by these authors contain only the 3′ terminal restriction fragments [12]. They were not amplified prior to the experiments so that they faithfully represent the composition of mRNA in the cellular source. In addition, care has been taken that the source materials are prepared to be as homogeneous as possible by using well-characterized cell lines [13], selective primary culturing [14], and purification with antibodies or through careful dissection [15]. The clones in the libraries were randomly selected and sequenced for identifying the genes and for measuring the abundance of their transcripts. Although they carry little amino acid sequence information, the 3′ sequences correspond to the respective genes, and are termed gene signatures (GS) [16]. The resulting lists, showing active genes and their relative activities, are called expression profiles, and some of these (parts of the profiles) are shown in Table 1 . Gene signatures (GS) that identify genes, composition of mRNA (f: frequency of appearance expressed per mil.), and the names of the genes are shown. Blanks under ‘gene name’ indicate novel genes. For purposes of clarity only the 30 most active genes are shown. By compiling expression profiles from different parts of the body, the cells or tissues where any given gene is active can be identified. As genes are mapped in the body where they are active, this data set is called a ‘Bodymap’ (http://www.imcb.osaka-u.ac.jp/bodymap) [10, 17]. As of June 1996, about 13 000 genes have been mapped on the body. A part of the bodymap which focuses on genes encoding cytoskeletal proteins is shown in Table 2 . Acc#, accession number in GenBank. For other definitions, see Table 1. Entries for human actins, tubulins, cytokeratins, and representative neural intermediate filaments were culled from Swiss Prot and a cDNA or a gene sequence corresponding to each of them in GenBank/EMBL was compared with Gene Signatures (GS) using the FastA program [33]. GS with bases more than 90% identical to those of the cDNAs or gene sequences were counted. Genes for cytokeratins 1, 2, 9, 10, 11, 12, and 20 were not found among either GS or ESTs (see Table 3). *Temporal lobes from an age matched pair of normal brain (1) and the brain of an Alzheimer disease patient (2). In addition to the bodymap, there are two major collections of partial cDNA sequences, collectively called EST (expressed sequence tags), which are compiled in a dbEST (http://www.ncbi.nlm.nih.gov/dbEST) [18]. One set, constructed in order to identify new genes of commercial interest, consists of 174 000 ESTs from randomly primed cDNA fragments or the 5′ ends of conventional cDNAs of human organs [19](for further references, see [20]). In this data set, several sequences from different regions of a single mRNA are collected indiscriminately. Another set containing some 280 000 ESTs [21]has been constructed to obtain probes for gene mapping [22]. The majority of the source cDNA libraries for this work were normalized' in vitro in order to reduce the number of abundant, frequently appearing clones [23]. Regardless of the original purpose of construction, these data sets are useful gene pools for ‘fishing’ new members of gene families or human orthologs of genes of other species, because they contain a significant number of entries [24, 25]. In Table 3 , distributions of ESTs among the human organs are shown. The same set of cytoskeletal gene transcripts as in Table 2 were selected for comparison. In accordance with the results in Table 2, actins and tubulins are distributed among a variety of tissues, as is already well known. Notice, however, that the multiple appearance of these gene transcripts in EST collections simply reflects incomplete normalization of the libraries, whereas those in the bodymap represent their gene activities. The high expression of tubulins in fetal neurons and fibroblasts, or the clear division of the site of cytokeratin gene expression into simple and stratified epithelial types, can be seen only in the bodymap because the histological resolution of gene expressions has been pursued. Overall, the dbEST should be regarded as a part of structural data, rather than functional data of the human genome. The same set of cytoskeletal gene sequences were compared with ESTs. Only EST entries in GenBank (re93, June96) annotated as 3′ were used to avoid multiple counting of ESTs representing the same mRNA molecule. The libraries subjected to normalization procedures are denoted as N. The numbers represent multiple appearances in the same library, reflecting insufficient normalization or excess amplification. As the number of genes collected and tissues analyzed in the bodymap is not yet large enough (the gene coverage is about one fourth of the dbEST), its usefulness remains limited. For this reason, a high throughput sequencing system is urgently needed. As a compromise, a method called SAGE has been proposed, in which 9 bp tags, resected from defined positions of cDNAs, are tandemly ligated and sequenced [26]. Methods other than nucleotide sequencing can also be employed for identifying active genes. Kato [27]has described ‘molecular indexing’, using the size of a restriction fragment of cDNA as an identifier: the cDNAs are cleaved by a type IIS enzyme and subjected to 64 different adapter-mediated PCRs. Altogether, 256 groups of the amplification products from a library have been fractionated in sequencing gels. By repeating this procedure with a few type IIS enzymes, most of the transcripts in source cells are displayed separately in the gel, one gene transcript being represented by a band of unique size, and its abundance by the band intensity. Another emerging technique is construction of arrays of oligonucleotides or unique fragments of cDNA at high density on solid support, which can be hybridized with uniformly labeled mRNA or cDNA [28, 29], for detection of active genes and their relative activities by their intensities. At the moment, however, specificity and sensitivity are the problems, since hybridization parameters differ from sequence to sequence. An average higher eukaryotic cell carries some 10 000 species of gene transcripts, and the total number of mRNA molecules is estimated to be several hundreds of thousands. Expression profiling, by sequencing or by other means, to describe the composition of these mRNA populations provides the basis for future functional analyses of the genome. The method established for the global description of gene expression control and gene networks can be extended to a variety of other systems and organisms. Some of the expected applications may include sorting novel genes that are defined as cell or tissue specific and studying their possible medical, diagnostic and pharmaceutical applications. Obtaining markers for monitoring cell differentiation or detecting pathological changes of gene activities in cells for diagnostic purposes may be other interesting examples of the application of this method. There is a great need for two major technological developments: high throughputs, as noted in the text, and down-scaling the analysis to a single cell level. Even though the system is still far from being technologically mature, the importance of the project demands enriching the databases in order to achieve the goal of describing how the approximately 100 000 genes in the human genome act in concert in the regulation of the whole body. Collaboration towards this goal is of crucial importance.
No takes yet. Share an insight, caveat, or question.
Okubo et al. (1997) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: