PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2009251 citationsOpen Access

Data quality from crowdsourcing

PHPei-Yun HsuehPMPrem MelvilleVSVikas Sindhwani

Key Points

Key points are not available for this paper at this time.

Abstract

Annotation acquisition is an essential step in training supervised classifiers. However, manual annotation is often time-consuming and expensive. The possibility of recruiting annotators through Internet services (e.g., Amazon Mechanic Turk) is an appealing option that allows multiple labeling tasks to be outsourced in bulk, typically with low overall costs and fast completion rates. In this paper, we consider the difficult problem of classifying sentiment in political blog snippets. Annotation data from both expert annotators in a research lab and non-expert annotators recruited from the Internet are examined. Three selection criteria are identified to select high-quality annotations: noise level, sentiment ambiguity, and lexical uncertainty. Analysis confirm the utility of these criteria on improving data quality. We conduct an empirical study to examine the effect of noisy annotations on the performance of sentiment classification models, and evaluate the utility of annotation selection on classification accuracy and efficiency.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hsueh et al. (2009) studied this question.

synapsesocial.com/papers/6a1cfe2ee19a8dd1302f1aafhttps://doi.org/10.3115/1564131.1564137
Ask AI
Helpful
Bookmark
Share
View Full Paper