PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 6, 2026Big Data and Cognitive Computing0 citationsOpen Access

Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research

View Full Paper
IHItai HimelboimMBMudit Baid

Key Points

  • This research aims to establish a framework for assessing the effectiveness of automated classifiers in harmful language detection.
  • Conducted a structured review of documentation practices for 27 publicly available classifiers.
  • Performed a cross-dataset evaluation to test models beyond their original training context.
  • Benchmarked a large language model (GPT-5) under a consistent prompting protocol.
  • Documentation practices for classifiers are often uneven and insufficient for theoretical measurement.
  • Inter-annotator agreement showed significant variability across datasets.
  • Cross-dataset performance frequently dropped, implying issues with generalizability.

Abstract

This manuscript develops and demonstrates a practical framework for evaluating automated classifiers used in communication research, using harmful language detection as an illustrative case. We combine (a) a structured review of documentation practices for 27 publicly available classifiers and their associated annotation processes with (b) a cross-dataset evaluation that re-tests each model beyond its original training context. Across 27 datasets, we extract and compare reporting on construct definitions, annotator instructions, and inter-annotator agreement, and we quantify generalization by applying each model to multiple out-of-domain test sets. We also benchmark a contemporary large language model (GPT-5) under a consistent prompting protocol to illustrate how LLM-based classification compares to fine-tuned classifiers. Results show that documentation is uneven and often insufficient for theory-driven measurement, inter-annotator agreement varies widely across datasets, and cross-dataset performance frequently drops substantially relative to within-dataset evaluations. Building on these findings and existing validation guidance, we provide a reusable checklist and decision flow to help researchers select, justify, and report classifier-based measures in ways that support transparency and cumulative science. Recommendations for researchers, reviewers, and journal editors stress aligning model selection with standards of validity, reliability, and transparency.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Himelboim et al. (2026) studied this question.

synapsesocial.com/papers/69fadb0b03f892aec9b1e939https://doi.org/10.3390/bdcc10050143
Ask AI
Helpful
Bookmark
Share
View Full Paper