We present two semi-supervised learning techniques to improve a state-of-the-art multi-lingual name tagger. For English and Chinese, the overall system obtains 1.7% -2.1% improvement in F-measure, representing a 13.5% -17.4% relative reduction in the spurious, missing, and incorrect tags. We also conclude that simply relying upon large corpora is not in itself sufficient: we must pay attention to unlabeled data selection too. We describe effective measures to automatically select documents and sentences.
No takes yet. Share an insight, caveat, or question.
Ji et al. (2006) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: