PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 24, 2018International Journal of Research in Marketing408 citationsOpen Access

Comparing automated text classification methods

View Full Paper
JHJochen HartmannJHJuliana HuppertzCSChristina Schamp

Key Points

  • This research aims to compare the performance of various automated text classification methods for analyzing social media data.
  • Analyzed ten automated text classification methods across 41 social media datasets.
  • Included five lexicon-based approaches and five machine learning algorithms.
  • Evaluated performance in terms of accuracy in classifying sentiment and content categories.
  • Random forest (RF) outperforms all methods for three-class sentiment classification.
  • Naive Bayes (NB) is most effective for small sample sizes.
  • Lexicon-based methods, particularly LIWC, show poor performance compared to machine learning techniques.

Abstract

Online social media drive the growth of unstructured text data. Many marketing applications require structuring this data at scales non-accessible to human coding, e.g., to detect communication shifts in sentiment or other researcher-defined content categories. Several methods have been proposed to automatically classify unstructured text. This paper compares the performance of ten such approaches (five lexicon-based, five machine learning algorithms) across 41 social media datasets covering major social media platforms, various sample sizes, and languages. So far, marketing research relies predominantly on support vector machines (SVM) and Linguistic Inquiry and Word Count (LIWC). Across all tasks we study, either random forest (RF) or naive Bayes (NB) performs best in terms of correctly uncovering human intuition. In particular, RF exhibits consistently high performance for three-class sentiment, NB for small samples sizes. SVM never outperform the remaining methods. All lexicon-based approaches, LIWC in particular, perform poorly compared with machine learning. In some applications, accuracies only slightly exceed chance. Since additional considerations of text classification choice are also in favor of NB and RF, our results suggest that marketing research can benefit from considering these alternatives.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hartmann et al. (2018) studied this question.

synapsesocial.com/papers/6a0f567f590fe99bbbed20b9https://doi.org/10.1016/j.ijresmar.2018.09.009
Ask AI
Helpful
Bookmark
Share
View Full Paper