We investigate various strategies for finding chemicals in biomedical text using substring co-occurrence information. The goal is to build a system from readily available data with minimal human involvement. Our models are trained from a dictionary of chemical names and general biomedical text. We investigated several strategies including Naïve Bayes classifiers and several types of N-gram models. We introduced a new way of interpolating N-grams that does not require tuning any parameters. We also found the task to be similar to Language Identification.
No takes yet. Share an insight, caveat, or question.
Alexander Vasserman (2004) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: