PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 25, 2017204 citationsOpen Access

A Large Labeled Corpus for Online Harassment Research

JGJennifer GolbeckZAZahra AshktorabRBRashad O. Banjo

Key Points

Key points are not available for this paper at this time.

Abstract

A fundamental part of conducting cross-disciplinary web science research is having useful, high-quality datasets that provide value to studies across disciplines. In this paper, we introduce a large, hand-coded corpus of online harassment data. A team of researchers collaboratively developed a codebook using grounded theory and labeled 35,000 tweets. Our resulting dataset has roughly 15% positive harassment examples and 85% negative examples. This data is useful for training machine learning models, identifying textual and linguistic features of online harassment, and for studying the nature of harassing comments and the culture of trolling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Golbeck et al. (2017) studied this question.

synapsesocial.com/papers/69d99aea2a25b240b7a3cf98https://doi.org/10.1145/3091478.3091509
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Detecting cyberbullying2013 · 119 citations
  2. 2Social Media Update 20162016 · 499 citations
  3. 3Trolling in asynchronous computer-mediated communication:from user discussions to theoretical concepts2010 · 16 citations
  4. 4Trolling in asynchronous computer-mediated communication: From user discussions to academic definitions2010 · 598 citations
  5. 5Automatic identification of personal insults on social news sites2011 · 154 citations