PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20254 citationsOpen Access

Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages

View Full Paper
FSFarhana ShahidMEMona ElswahAVAditya Vashistha

Key Points

  • Automated moderation tools for low-resource languages face significant structural inequities, highlighting the need for reform.
  • Interviews with 22 AI experts exposed that socio-political factors exacerbate the challenges of data scarcity in low-resource languages.
  • The study emphasizes shortcomings in AI language models regarding complex linguistic features, calling for more inclusive design.
  • Historical colonial suppression continues to impact language resources, which undermines effective moderation in diverse languages.

Abstract

Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews with 22 AI experts working on harmful content detection in four low-resource languages: Tamil (South Asia), Swahili (East Africa), Maghrebi Arabic (North Africa), and Quechua (South America)--we examine systemic issues in building automated moderation tools for these languages. Our findings reveal that beyond data scarcity, socio-political factors such as tech companies' monopoly on user data and lack of investment in moderation for low-profit Global South markets exacerbate historic inequities. Even if more data were available, the English-centric and data-intensive design of language models and preprocessing techniques overlooks the need to design for morphologically complex, linguistically diverse, and code-mixed languages. We argue these limitations are not just technical gaps caused by "data scarcity" but reflect structural inequities, rooted in colonial suppression of non-Western languages. We discuss multi-stakeholder approaches to strengthen local research capacity, democratize data access, and support language-aware solutions to improve automated moderation for low-resource languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shahid et al. (2025) studied this question.

synapsesocial.com/papers/68f12bfb2107091eab27a4c0https://doi.org/10.1609/aies.v8i3.36719
Ask AI
Helpful
Bookmark
Share
View Full Paper