PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 201637 citationsOpen Access

Learning a POS tagger for AAVE-like language

AJAnna JørgensenDHDirk HovyASAnders Søgaard

Key Points

Key points are not available for this paper at this time.

Abstract

Part-of-speech (POS) taggers trained on newswire perform much worse on domains such as subtitles, lyrics, or tweets. In addition, these domains are also heterogeneous, e.g., with respect to registers and dialects. In this paper, we consider the problem of learning a POS tagger for subtitles, lyrics, and tweets associated with African-American Vernacular English (AAVE). We learn from a mixture of randomly sampled and manually annotated Twitter data and unlabeled data, which we automatically and partially label using mined tag dictionaries. Our POS tagger obtains a tagging accuracy of 89% on subtitles, 85% on lyrics, and 83% on tweets, with up to 55% error reductions over a state-of-the-art newswire POS tagger, and 15-25% error reductions over a state-of-the-art Twitter POS tagger.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jørgensen et al. (2016) studied this question.

synapsesocial.com/papers/6a0f3dd89cac01975e42859ehttps://doi.org/10.18653/v1/n16-1130
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1POS tagging of dialectal Arabic2005 · 51 citations
  2. 2Hassle-free POS-Tagging for the Alsatian Dialects2013 · 5 citations
  3. 3Mining for unambiguous instances to adapt POS taggers to new domains2015 · 2 citations
  4. 4Wiki-ly Supervised Part-of-Speech Tagging2012 · 84 citations
  5. 5Semi-supervised learning and domain adaptation for NLP2013 · 4 citations