The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner. The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner. In November 2004, an article was published in Nature by the International Human Genome Sequencing Centre announcing the finishing of the sequencing of the human genome (1.Stein L.D. Human genome: end of the beginning.Nature. 2004; 431: 915-916Google Scholar). The published sequence covered 99% of the euchromatic genome and contained only 341 gaps. This incredible achievement has rightly been hailed as a foundation for biomedical research in the decades ahead but, in practice, is only the first step in a long and complicated path to decipher the complexity of the proteome content of the human cell.To fully understand the workings of the human proteome, scientists must first be able to identify every protein coding region contained within the genome and the amino acid sequence of the proteins that these regions encode. In addition to this basic information, an incredible amount of metadata needs to be assembled. For example, the signals that trigger the expression of these proteins must be identified, the actual protein expression experimentally observed and catalogued. The subsequent duration of gene expression along with the factors can control its eventual repression, the stability of mRNA transcripts and the rates at which they are translated into protein products must also be known and understood. Every potential site of posttranslational modification of the protein should be identified, the conditions under which these modifications are made and their biological significance understood. The biological function of each protein molecule needs to be catalogued along with how this varies according to the cell type in which it is and the within the The significance of the each protein with proteins, and also needs be at both a and functional and in with of the and by the only this information first to be and this is in over the but the information needs to be and in a that it to with an interest in the of this data is already in published and is to with every the potential user is with a should they to information on a gene and this with that to in this domain databases to this information and it to a the user to the how many protein coding genes are present in the human genome has been a that has scientists long the of sequencing In the to be the of human on the of on data and of sequence human many genes in the human Scholar). an the of experimentally in the which from to of to the of gene This is by An of 2004; a that biological information the of in to the of the to sequence the human genome, gene which gene structure protein and 2004; and which a gene structure and structure and 2004; Scholar). gene at the of of The of protein coding genes by has with each and of the the by of the of coding on the of the human genome is also coding from the human genome sequence of transcripts and of these is by the International International The International An for 2004; which was first for the of the human genome the experimentally protein in the sequence The with the protein predictions of and both protein predictions and experimentally by to a of and proteins of sequence are in as their protein are is by the on the of protein and the data is but to the of in proteins from databases and a sequence be identified, are and can be by the in a are as a of data within the a be to the to be of the human to be transcripts from the human genome, with only of by is to be that the of these gene products can be experimentally over the to a of the human it is to be that the of as to both their and to experimentally the predictions are also to a of transcript and protein for the human protein sequence information from a of the of transcripts from many as genome and gene sequencing in addition to data by protein the for a these be into a and with functional and information. The was to this and was the of the existing The protein and its in The protein and its in and The sequence is in a the the of and is of each for The is the for protein information, protein and The databases into a to The is a the of protein of and protein from many are to a which protein products by an gene from a are to sequencing and to identify both and of of modifications in the 2004; Scholar). are and that each sequence be from within the of posttranslational modification are identified, and by are as The protein is both a protein and gene and known are data and information are and information on the protein is the annotation on as the of the information and by of the expression of the to proteins, of the protein in a with in the of the protein as a of annotation is and the at which the can of sequence was in and of from the of coding in the sequence for coding already in also protein from the by the user community that are in has a of sequence a gene from an be by The data content is by annotation protein 2004; Scholar). The a of human this many which be into a within of the many of the is the made to can the of information on a protein but to data protein and and databases be as a of which to many to the information in the has a to known human according to the of Human The human Scholar). human protein been fully with an within these Human of of human in of of and of posttranslational modifications of to published of of of of to of to of to of to in a the of annotation is and can only data that has been experimentally for a protein in a In to of this information to proteins within the must be a of of proteins functional regions within of and sequence for protein of these been and into an and in Scholar). is from by The its in The for identification of protein The protein and for protein domain and genome and annotation of from and protein The of protein protein and in structure and sequence 2004; with protein for proteins and within the and posttranslational modification are to the and databases with proteins and are to the for proteins that the they the within that the for the in the to a novel protein sequence and function by to known protein and to identify functional within the is within as the for annotation from the to proteins in the This information to a of the protein in the the human genome proteins, of the protein expression is to the protein content of a cell in a The of protein expression is in a of the The Human and data and 2004; Scholar). in the of protein expression data is the of and data in the The for Public of and 2004; in the of for 2004; community for to the and of data been by the within the which the and of information, and which both protein and the from which the identification was The these and a for protein identification which is to and data in function in and the of a protein with the in a cell at which the molecule is the in which it is and the of the with which it is of is to a of in this information within the but this by to data For example, protein data is in a An 2004; Scholar). within is from from existing by the by to and made to the to also a of for and the for and a and an that for protein data has and An a of from and the conditions under which these been An only a of in the of An is a biological in an a but also a a An in the of both data and the of to and of the for it is to by that is fully with the and can and data in both and The community for the of protein 2004; Scholar). is also a of the a of also The The of research for of protein and and The ahead of which to data to an at of the information, the and that these is and in databases as of a 2004; Scholar). is by biological with in their and and by the to the proteins by to with to from the information as to which each protein a on the human proteome is over an of databases and a of must be to information on a protein to be and of a protein as the of a gene as by the Human Gene of a 2004; of in that the protein can be the in data on the of and in this are the that to the of gene the of their the biological in which they a and in which they are The Human Gene 2004; Scholar). The which from a of to human annotation is and the can be by annotation on In this a been to human proteins The Gene in with Gene 2004; Scholar). are many of the databases by and and in this the of to gene expression for the of the which to cell and The Gene in with Gene 2004; and also the at biological of these are at the site are a long from a of the human proteome, in of the each molecule in the but is both and of data is and made an of a of protein is by which protein in they the protein a and information for protein The the human protein with and human sequence into a of known human protein and to and and to the of human proteome in a a for and In November 2004, an article was published in Nature by the International Human Genome Sequencing Centre announcing the finishing of the sequencing of the human genome (1.Stein L.D. Human genome: end of the beginning.Nature. 2004; 431: 915-916Google Scholar). The published sequence covered 99% of the euchromatic genome and contained only 341 gaps. This incredible achievement has rightly been hailed as a foundation for biomedical research in the decades ahead but, in practice, is only the first step in a long and complicated path to decipher the complexity of the proteome content of the human fully understand the workings of the human proteome, scientists must first be able to identify every protein coding region contained within the genome and the amino acid sequence of the proteins that these regions encode. In addition to this basic information, an incredible amount of metadata needs to be assembled. For example, the signals that trigger the expression of these proteins must be identified, the actual protein expression experimentally observed and catalogued. The subsequent duration of gene expression along with the factors can control its eventual repression, the stability of mRNA transcripts and the rates at which they are translated into protein products must also be known and understood. Every potential site of posttranslational modification of the protein should be identified, the conditions under which these modifications are made and their biological significance understood. The biological function of each protein molecule needs to be catalogued along with how this varies according to the cell type in which it is and the within the The significance of the each protein with proteins, and also needs be at both a and functional and in with of the and by the only this information first to be and this is in over the but the information needs to be and in a that it to with an interest in the of this data is already in published and is to with every the potential user is with a should they to information on a gene and this with that to in this domain databases to this information and it to a the user to the how many protein coding genes are present in the human genome has been a that has scientists long the of sequencing In the to be the of human on the of on data and of sequence human many genes in the human Scholar). an the of experimentally in the which from to of to the of gene This is by An of 2004; a that biological information the of in to the of the to sequence the human genome, gene which gene structure protein and 2004; and which a gene structure and structure and 2004; Scholar). gene at the of of The of protein coding genes by has with each and of the the by of the of coding on the of the human genome is also coding from the human genome sequence of transcripts and of these is by the International International The International An for 2004; which was first for the of the human genome the experimentally protein in the sequence The with the protein predictions of and both protein predictions and experimentally by to a of and proteins of sequence are in as their protein are is by the on the of protein and the data is but to the of in proteins from databases and a sequence be identified, are and can be by the in a are as a of data within the a be to the to be of the human to be transcripts from the human genome, with only of by is to be that the of these gene products can be experimentally over the to a of the human it is to be that the of as to both their and to experimentally the predictions are also to a of transcript and protein for the human how many protein coding genes are present in the human genome has been a that has scientists long the of sequencing In the to be the of human on the of on data and of sequence human many genes in the human Scholar). an the of experimentally in the which from to of to the of gene This is by An of 2004; a that biological information the of in to the of the to sequence the human genome, gene which gene structure protein and 2004; and which a gene structure and structure and 2004; Scholar). gene at the of of The of protein coding genes by has with each and of the the by of the of coding on the of the human genome is also coding from the human genome sequence of transcripts and Scholar). of these is by the International International The International An for 2004; which was first for the of the human genome the experimentally protein in the sequence The with the protein predictions of and both protein predictions and experimentally by to a of and proteins of sequence are in as their protein are is by the on the of protein and the data is but to the of in proteins from databases and a sequence be identified, are and can be by the in a are as a of data within the a be to the to be of the human to be transcripts from the human genome, with only of by is to be that the of these gene products can be experimentally over the to a of the human it is to be that the of as to both their and to experimentally the predictions are The also to a of transcript and protein for the human protein sequence information from a of the of transcripts from many as genome and gene sequencing in addition to data by protein the for a these be into a and with functional and information. The was to this and was the of the existing The protein and its in The protein and its in and The sequence is in a the the of and is of each for The is the for protein information, protein and The databases into a to The is a the of protein of and protein from many are to a which protein products by an gene from a are to sequencing and to identify both and of of modifications in the 2004; Scholar). are and that each sequence be from within the of posttranslational modification are identified, and by are as The protein is both a protein and gene and known are data and information are and information on the protein is the annotation on as the of the information and by of the expression of the to proteins, of the protein in a with in the of the protein as a of annotation is and the at which the can of sequence was in and of from the of coding in the sequence for coding already in also protein from the by the user community that are in has a of sequence a gene from an be by The data content is by annotation protein 2004; Scholar). The a of human this many which be into a within of the many of the is the made to can the of information on a protein but to data protein and and databases be as a of which to many to the information in the has a to known human according to the of Human The human Scholar). human protein been fully with an within these Human of of human in of of and of posttranslational modifications of to published of of of of to of to of to of to in a protein sequence information from a of the of transcripts from many as genome and gene sequencing in addition to data by protein the for a these be into a and with functional and information. The was to this and was the of the existing The protein and its in The protein and its in and The sequence is in a the the of and is of each for The is the for protein information, protein and The databases into a to The is a the of protein The of and protein from many are to a which protein products by an gene from a are to sequencing and to identify both and of of modifications in the 2004; Scholar). are and that each sequence be from within the of posttranslational modification are identified, and by are as The protein is both a protein and gene and known are data and information are and information on the protein is the annotation on as the of the information and by of the expression of the to proteins, of the protein in a with in the of the protein as a of annotation is and the at which the can of sequence was in and of from the of coding in the sequence for coding already in also protein from the by the user community that are in has a of sequence a gene from an be by The data content is by annotation protein 2004; Scholar). The a of human this many which be into a within of the many of the is the made to can the of information on a protein but to data protein and and databases be as a of which to many to the information in the has a to known human according to the of Human The human Scholar). human protein been fully with an within these the of annotation is and can only data that has been experimentally for a protein in a In to of this information to proteins within the must be a of of proteins functional regions within of and sequence for protein of these been and into an and in Scholar). is from by The its in The for identification of protein The protein and for protein domain and genome and annotation of from and protein The of protein protein and in structure and sequence 2004; with protein for proteins and within the and posttranslational modification are to the and databases with proteins and are to the for proteins that the they the within that the for the in the to a novel protein sequence and function by to known protein and to identify functional within the is within as the for annotation from the to proteins in the This information to a of the protein in the the of annotation is and can only data that has been experimentally for a protein in a In to of this information to proteins within the must be a of of proteins functional regions within of and sequence for protein of these been and into an and in Scholar). is from by The its in The for identification of protein The protein and for protein domain and genome and annotation of from and protein The of protein protein and in structure and sequence 2004; with protein for proteins and within the and posttranslational modification are to the and databases with proteins and are to the for proteins that the they the within that the for the in the to a novel protein sequence and function by to known protein and to identify functional within the is within as the for annotation from the to proteins in the This information to a of the protein in the the human genome proteins, of the protein expression is to the protein content of a cell in a The of protein expression is in a of the The Human and data and 2004; Scholar). in the of protein expression data is the of and data in the The for Public of and 2004; in the of for 2004; community for to the and of data been by the within the which the and of information, and which both protein and the from which the identification was The these and a for protein identification which is to and data in function in and the of a protein with the in a cell at which the molecule is the in which it is and the of the with which it is of is to a of in this information within the but this by to data For example, protein data is in a An 2004; Scholar). within is from from existing by the by to and made to the to also a of for and the for and a and an that for protein data has and An a of from and the conditions under which these been An only a of in the of An is a biological in an a but also a a An in the of both data and the of to and of the for it is to by that is fully with the and can and data in both and The community for the of protein 2004; Scholar). is also a of the a of also The The of research for of protein and and The ahead of which to data to an at of the information, the and that these is and in databases as of a 2004; Scholar). is by biological with in their and and by the to the proteins by to with to from the information as to which each protein a the human genome proteins, of the protein expression is to the protein content of a cell in a The of protein expression is in a of the The Human and data and 2004; Scholar). in the of protein expression data is the of and data in the The for Public of and 2004; in the of for 2004; community for to the and of data been by the within the which the and of information, and which both protein and the from which the identification was The these and a for protein identification which is to and data in function in and the of a protein with the in a cell at which the molecule is the in which it is and the of the with which it is of is to a of in this information within the but this by to data For example, protein data is in a An 2004; Scholar). within is from from existing by the by to and made to the to also a of for and the for and a and an that for protein The data has and An a of from and the conditions under which these been An only a of in the of An is a biological in an a but also a a An in the of both data and the of to and of the for it is to by that is fully with the and can and data in both and The community for the of protein 2004; Scholar). is also a of the a of also The The of research for of protein and and The ahead of which to data to an at of the information, the and that these is and in databases as of a 2004; Scholar). is by biological with in their and and by the to the proteins by to with to from the information as to which each protein a on the human proteome is over an of databases and a of must be to information on a protein to be and of a protein as the of a gene as by the Human Gene of a 2004; of in that the protein can be the in data on the of and in this are the that to the of gene the of their the biological in which they a and in which they are The Human Gene 2004; Scholar). The which from a of to human annotation is and the can be by annotation on In this a been to human proteins The Gene in with Gene 2004; Scholar). are many of the databases by and and in this the of to gene expression for the of the which to cell and The Gene in with Gene 2004; and also the at biological of these are at the site on the human proteome is over an of databases and a of must be to information on a protein to be and of a protein as the of a gene as by the Human Gene of a 2004; of in that the protein can be the in data on the of and in this are the that to the of gene the of their the biological in which they a and in which they are The Human Gene 2004; Scholar). The which from a of to human annotation is and the can be by annotation on In this a been to human proteins The Gene in with Gene 2004; Scholar). are many of the databases by and and in this the of to gene expression for the of the which to cell and The Gene in with Gene 2004; and also the at biological of these are at the site are a long from a of the human proteome, in of the each molecule in the but is both and of data is and made an of a of protein is by which protein in they the protein a and information for protein The the human protein with and human sequence into a of known human protein and to and and to the of human proteome in a a for and are a long from a of the human proteome, in of the each molecule in the but is both and of data is and made an of a of protein is by which protein in they the protein a and information for protein The the human protein with and human sequence into a of known human protein and to and and to the of human proteome in a a for and
No takes yet. Share an insight, caveat, or question.
Orchard et al. (2005) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: