PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 202448 citationsOpen Access

LLM-Mod: Can Large Language Models Assist Content Moderation?

View Full Paper
MKMahi KollaSSSiddharth SalunkheECEshwar Chandrasekharan

Key Points

Key points are not available for this paper at this time.

Abstract

Content moderation is critical for maintaining healthy online spaces. However, it remains a predominantly manual task. Moderators are often exhausted by low moderator-to-posts ratio. Researchers have been exploring computational tools to assist human moderators. The natural language understanding capabilities of large language models (LLMs) open up possibilities to use LLMs for online moderation. This work explores the feasibility of using LLMs to identify rule violations on Reddit. We examine how an LLM-based moderator (LLM-Mod) reasons about 744 posts across 9 subreddits that violate different types of rules. We find that while LLM-Mod has a good true-negative rate (92.3%), it has a bad true-positive rate (43.1%), performing poorly when flagging rule-violating posts. LLM-Mod is likely to flag keyword-matching-based rule violations, but cannot reason about posts with higher complexity. We discuss the considerations for integrating LLMs into content moderation workflows and designing platforms that support both AI-driven and human-in-the-loop moderation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kolla et al. (2024) studied this question.

synapsesocial.com/papers/68e6bbd2b6db64358763c991https://doi.org/10.1145/3613905.3650828
Ask AI
Helpful
Bookmark
Share
View Full Paper