Key points are not available for this paper at this time.
The automatic classification of digital hate is a pressing challenge, yet many existing computational models remain opaque and insufficiently evaluated as measurement tools in social science contexts. This study examines the utility of Google’s Perspective API as a measurement instrument by modeling higher-order constructs of harmful discourse (i.e., incivility and intolerance) as outcomes of multiple lower-level behavioral indicators captured by distinct API scores rather than using a single aggregate score as a proxy for complex social behaviors while enabling the evaluation of state-of-the-art black-box classifiers beyond classification metrics. Drawing on 4,000 manually annotated English-language YouTube comments in the context of the Israel-Hamas war, we test whether multiple API scores predict incivility and intolerance using generalized linear and additive models, assess classification performance across non-hateful, uncivil, and intolerant content, and benchmark a recent deep learning model. Results show that Identity Attack is a strong predictor of intolerance, whereas Insult and Profanity are indicative of incivility. While classification performance is somewhat below state-of-the-art deep learning models, our approach offers important advantages: transparency, interpretability, accessibility for non-technical researchers, and potential cross-linguistic applicability. We argue that typology-driven, multi-indicator-based classification provides a practical and theoretically grounded complement to more aggregated black-box models, particularly in human-in-the-loop workflows that can help reduce annotator exposure through pre-filtering of content.
Kirchmair et al. (Sat,) studied this question.