PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 1, 2005BMC Bioinformatics47 citationsOpen Access

Learning Statistical Models for Annotating Proteins with Function Information using Biomedical Text

SRSoumya RayMCMark Craven

Key Points

Key points are not available for this paper at this time.

Abstract

BACKGROUND: The BioCreative text mining evaluation investigated the application of text mining methods to the task of automatically extracting information from text in biomedical research articles. We participated in Task 2 of the evaluation. For this task, we built a system to automatically annotate a given protein with codes from the Gene Ontology (GO) using the text of an article from the biomedical literature as evidence. METHODS: Our system relies on simple statistical analyses of the full text article provided. We learn n-gram models for each GO code using statistical methods and use these models to hypothesize annotations. We also learn a set of Naïve Bayes models that identify textual clues of possible connections between the given protein and a hypothesized annotation. These models are used to filter and rank the predictions of the n-gram models. RESULTS: We report experiments evaluating the utility of various components of our system on a set of data held out during development, and experiments evaluating the utility of external data sources that we used to learn our models. Finally, we report our evaluation results from the BioCreative organizers. CONCLUSION: We observe that, on the test data, our system performs quite well relative to the other systems submitted to the evaluation. From other experiments on the held-out data, we observe that (i) the Naïve Bayes models were effective in filtering and ranking the initially hypothesized annotations, and (ii) our learned models were significantly more accurate when external data sources were used during learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ray et al. (2005) studied this question.

synapsesocial.com/papers/6a82b59b516b1a94b7d4ad8dhttps://doi.org/10.1186/1471-2105-6-s1-s18
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The SWISS-PROT protein sequence data bank and its supplement TrEMBL1997 · 744 citations
  2. 2The FlyBase database of the Drosophila genome projects and community literature2003 · 490 citations
  3. 3Guidelines for Human Gene Nomenclature2002 · 503 citations
  4. 4The Arabidopsis Information Resource (TAIR): a comprehensive database and web-based information retrieval, analysis, and visualization system for a model plant2001 · 590 citations
  5. 5Unified Medical Language System2020 · 14 citations