January 1, 1994Open Access

Adaptive sentence boundary disambiguation

Key Points

Key points are not available for this paper at this time.

Abstract

Labeling of sentence boundaries is a necessary prerequisite for many natural language processing tasks, including part-ofspeech tagging and sentence alignment. End-of-sentence punctuation marks are ambiguous; to disambiguate them most systems use brittle, special-purpose regular expression grammars and exception rules. As an alternative, we have developed an efficient, trainable algorithm that uses a lexicon with part-of-speech probabilities and a feed-forward neural network. This work demonstrates the feasibility of using prior probabilities of part-of-speech assignments, as opposed to words or definite part-ofspeech assignments, as contextual information. After training for less than one minute, the method correctly labels over 98.5% of sentence boundaries in a corpus of over 27,000 sentence-boundary marks. We show the method to be efficient and easily adaptable to different text genres, including single-case texts.

KI fragen

Bookmark

View Full Paper

Cite This Study

Palmer et al. (Sat,) studied this question.

synapsesocial.com/papers/6a0e9ef2218372ada6479e3d https://doi.org/https://doi.org/10.3115/974358.974376

KI fragen

Bookmark

View Full Paper