PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 20, 2026ACM Transactions on Intelligent Systems and Technology0 citations

Ideology-Based LLMs for Content Moderation

View Full Paper
SCStefano CivelliPBPietro BernardelleNPNardiena Pratama

Key Points

  • The aim is to explore how ideological personas impact the consistency and fairness of harmful content classification in large language models.
  • Examined different LLM architectures and model sizes
  • Analyzed harmful content classification across language and vision modalities
  • Conducted agreement analyses to assess ideological alignment of models
  • Personas show distinct labeling tendencies based on ideological leanings
  • Larger models align closer with personas of the same political ideology
  • Ideological conditioning leads to subtle biases in content moderation outputs

Abstract

Large language models (LLMs) are increasingly used in content moderation systems, where ensuring fairness and objectivity is essential. In this study, we examine how persona adoption influences the consistency and fairness of harmful content classification across different LLM architectures, model sizes, and content modalities (i.e., language vs. vision). At first glance, headline performance metrics suggest that personas have little impact on overall classification accuracy. However, a closer analysis reveals important behavioral shifts. Personas with different ideological leanings display distinct propensities to label content as harmful, showing that the lens through which a model “views” input can subtly shape its judgments. Agreement analyses highlight that models, particularly larger ones, tend to align more closely with personas from the same political ideology, strengthening within-ideology consistency while widening divergence across ideological groups. To show this effect more directly, we conducted a study on a politically targeted task, which confirmed that personas not only behave more coherently within their own ideology but also exhibit a tendency to defend their perspective while downplaying harmfulness in opposing views. Together, these findings highlight how persona conditioning can introduce subtle ideological biases into LLM outputs, raising concerns about the use of AI systems that may reinforce partisan perspectives under the guise of neutrality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Civelli et al. (2026) studied this question.

synapsesocial.com/papers/69e5c36103c29399140292dfhttps://doi.org/10.1145/3810946
Ask AI
Helpful
Bookmark
Share
View Full Paper