Sociodemographic biases are a common problem for natural language processing, the fairness and integrity of its applications. Within sentiment, these biases may undermine sentiment predictions for texts that personal attributes that unbiased human readers would consider neutral. discrimination can have great consequences in the applications of analysis both in the public and private sectors. For example, inferences in applications like online abuse and opinion analysis in media platforms can lead to unwanted ramifications, such as wrongful, towards certain populations. In this paper, we address the against people with disabilities, PWD, done by sentiment and toxicity classification models. We provide an examination of and toxicity analysis models to understand in detail how they PWD. We present the Bias Identification Test in Sentiments (BITS), corpus of 1,126 sentences designed to probe sentiment analysis models for in disability. We use this corpus to demonstrate statistically biases in four widely used sentiment analysis tools (TextBlob,, Google Cloud Natural Language API and DistilBERT) and two toxicity models trained to predict toxic comments on Jigsaw challenges (Toxic classification and Unintended Bias in Toxic comments). The results show all exhibit strong negative biases on sentences that mention disability. publicly release BITS Corpus for others to identify potential biases against in any sentiment analysis tools and also to update the corpus to be as a test for other sociodemographic variables as well.
No takes yet. Share an insight, caveat, or question.
Venkit et al. (2021) studied this question.