We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.
No takes yet. Share an insight, caveat, or question.
Sproat et al. (1994) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: