Classifying proteins into families and superfamilies allows identification of functionally important con-served domains. The motifs and scoring matrices derived fromsuchconserved regionsprovidecompu-tational tools that recognize similar patterns in novel sequences, and thus enable the prediction of protein function for genomes. The eBLOCKs database enu-merates a cascade of protein blocks with varied conservation levels for each functional domain. A biologically important region is most stringently conserved among a smaller family of highly similar proteins. The same region is often found in a larger group of more remotely related proteins with a reduced stringency. Through enumeration, highly specific signatures can be generated from blocks with more columns and fewer family members, while highly sensitive signatures can be derived from blocks with fewer columns and more members as in a superfamily. By applyingPSI-BLAST and amodified K-means clustering algorithm, eBLOCKs automati-cally groupsprotein sequencesaccording todifferent levels of similarity. Multiple sequence alignments are made and trimmed into a series of ungapped blocks. Motifs and position-specific scoring matrices were derived from eBLOCKs and made available for sequence search and annotation. The eBLOCKsdata-base provides a tool for high-throughput genome annotation with maximal specificity and sensitivity. The eBLOCKs database is freely available on the WorldWideWeb at
No takes yet. Share an insight, caveat, or question.
Qin Su (2004) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: