PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2003104 citationsOpen Access

Boosting precision and recall of dictionary-based protein name recognition

YTYoshimasa TsuruokaJTJun’ichi Tsujii

Key Points

Key points are not available for this paper at this time.

Abstract

Dictionary-based protein name recognition is the first step for practical information extraction from biomedical documents because it provides ID information of recognized terms unlike machine learning based approaches. However, dictionary based approaches have two serious problems: (1) a large number of false recognitions mainly caused by short names. (2) low recall due to spelling variation. In this paper, we tackle the former problem by using a machine learning method to filter out false positives. We also present an approximate string searching method to alleviate the latter problem. Experimental results using the GE-NIA corpus show that the filtering using a naive Bayes classifier greatly improves precision with slight loss of recall, resulting in a much better F-score.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tsuruoka et al. (2003) studied this question.

synapsesocial.com/papers/6a0ea3afa14f152feaf9a48dhttps://doi.org/10.3115/1118958.1118964
Ask AI
Helpful
Bookmark
Share
View Full Paper