PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 25, 2017204 citationsOpen Access

A Large Labeled Corpus for Online Harassment Research

JGJennifer GolbeckZAZahra AshktorabRBRashad O. Banjo

Key Points

Key points are not available for this paper at this time.

Abstract

A fundamental part of conducting cross-disciplinary web science research is having useful, high-quality datasets that provide value to studies across disciplines. In this paper, we introduce a large, hand-coded corpus of online harassment data. A team of researchers collaboratively developed a codebook using grounded theory and labeled 35,000 tweets. Our resulting dataset has roughly 15% positive harassment examples and 85% negative examples. This data is useful for training machine learning models, identifying textual and linguistic features of online harassment, and for studying the nature of harassing comments and the culture of trolling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Golbeck et al. (2017) studied this question.

synapsesocial.com/papers/69d99aea2a25b240b7a3cf98https://doi.org/10.1145/3091478.3091509
Ask AI
Helpful
Bookmark
Share
View Full Paper