PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 201350 citationsOpen Access

Elephant: Sequence Labeling for Word and Sentence Segmentation

KEKilian EvangVBValerio BasileGCGrzegorz Chrupała

Key Points

Key points are not available for this paper at this time.

Abstract

Tokenization is widely regarded as a solved problem due to the high accuracy that rule-based tokenizers achieve. But rule-based tokenizers are hard to maintain and their rules language specific. Like an elephant in the living room, it is a problem that is impossible to overlook whenever new raw datasets need to be processed or when tokenization conventions are reconsidered. It is moreover an important problem, because any errors occurring early in the pipeline affect further analysis negatively. We believe that regarding tokenization, there is still room for improvement, in particular on the methodological side of the task. We are particularly interested in the following questions: Can we use supervised learning to avoid hand-crafting rules? Can we use unsupervised feature learning to reduce feature engineering effort and boost performance? Can we use the same method across languages? Can we combine word and sentence boundary detection into one task?

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Evang et al. (2013) studied this question.

synapsesocial.com/papers/6a0db17548a82a5ce309d072https://doi.org/10.18653/v1/d13-1146
Ask AI
Helpful
Bookmark
Share
View Full Paper