PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 27, 2026Language Resources and Evaluation0 citationsOpen Access

Toxicbias-reasoning: a multicultural dataset for social bias detection with human-aligned reasoning

View Full Paper
AKAnuj KumarMGMahendra Kumar GurveSASatyadev Ahlawat

Key Points

  • The central aim is to enhance social bias detection in multilingual and multicultural contexts through a diverse dataset.
  • Developed a dataset of 7,562 annotated bias-related statements.
  • Included culturally diverse topics such as race, gender, and caste.
  • Employed transformer-based models with hierarchical and multi-task training strategies.
  • Conducted evaluations using human-in-the-loop processes for training and validation.
  • Achieved a macro F1 score of 0.9099 with RoBERTa for bias detection.
  • Incorporated logic-aware loss improved consistency, yielding F1 scores of up to 0.95 for major categories.
  • Notable gains in minority categories like caste and political bias were observed.

Abstract

Social bias in language models continues to create fairness risks in multilingual and multicultural environments. Existing datasets provide limited cultural diversity, insufficient support for overlapping bias categories, and minimal availability of human-interpretable reasoning, which reduces transparency and reliability in the bias detection. The ToxicBias-Reasoning dataset addresses these gaps by providing 7,562 annotated statements representing bias related to caste, religion, race, gender, political identity, and LGBTQ issues, including 247 caste-specific samples and 1,923 non-biased samples. The dataset includes manually validated labels and instance-level reasoning, with training and validation reasoning generated through a GPT-4o-assisted human-in-the-loop process and a fully manual test set for high-fidelity evaluation. This study evaluates transformer-based classification models trained using hierarchical and multi-task strategies. The results demonstrate strong performance for bias detection, with RoBERTa achieving a macro F1 score of 0.9099, and for multilabel category classification, where incorporating a logic-aware loss improves consistency and yields F1 scores of up to 0.95 for major categories such as race and religion, along with notable gains in minority categories, including caste and political bias. A text-to-text knowledge distillation framework additionally trains a compact generative model for reasoning, with BART-Large attaining a ROUGE-L score of 45.22 and a BLEU score of 16.90. These findings support the practical deployment of explainable and culturally grounded bias detection systems in fairness-critical natural language processing applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kumar et al. (2026) studied this question.

synapsesocial.com/papers/69c620ab15a0a509bde193f2https://doi.org/10.1007/s10579-026-09916-w
Ask AI
Helpful
Bookmark
Share
View Full Paper