Just like other NLP applications, a serious problem with Chinese word segmentation lies in the ambiguities involved. Disambiguation methods fall into different categories, e.g., rule-based, statistical-based and example-based approaches, each of which may involve a variety of machine learning techniques. In this paper we report our current progress within the example-based approach, including its framework, example representation and collection, example matching and application. Experimental results show that this effective approach resolves more than 90% of ambiguities found. Hence, if it is integrated effectively with a segmentation method of the precision P > 95%, the resulting segmentation accuracy can reach, theoretically, beyond 99.5%.
No takes yet. Share an insight, caveat, or question.
Kit et al. (2002) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: