We present a language-independent and unsupervised algorithm for the segmentation of words into morphs. The algorithm is based on a new generative probabilistic model, which makes use of relevant prior information on the length and frequency distributions of morphs in a language. Our algorithm is shown to outperform two competing algorithms, when evaluated on data from a language with agglutinative morphology (Finnish), and to perform well also on English data.
No takes yet. Share an insight, caveat, or question.
Mathias Creutz (2003) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: