PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 8, 20240 citationsOpen Access

Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate

View Full Paper
YZYiqun ZhangXYXiaocui YangFSFeng Shi

Key Points

  • Agent4Debate demonstrates comparable performance to human debaters in competitive debate scenarios, indicating promising advancements in AI debate.
  • Using the Debatrix scoring system, results show the agent's effectiveness across 200 debates involving diverse motions and formats.
  • The framework comprises four agents: Searcher, Analyzer, Writer, and Reviewer, each specializing in distinct phases of debate preparation and execution for enhanced collaboration and performance.  Ablation studies confirm the significance of each component within the framework, supporting future developments in AI argumentation and debate systems.

Abstract

Competitive debate is a complex task of computational argumentation. Large Language Models (LLMs) suffer from hallucinations and lack competitiveness in this field. To address these challenges, we introduce Agent for Debate (Agent4Debate), a dynamic multi-agent framework based on LLMs designed to enhance their capabilities in competitive debate. Drawing inspiration from human behavior in debate preparation and execution, Agent4Debate employs a collaborative architecture where four specialized agents, involving Searcher, Analyzer, Writer, and Reviewer, dynamically interact and cooperate. These agents work throughout the debate process, covering multiple stages from initial research and argument formulation to rebuttal and summary. To comprehensively evaluate framework performance, we construct the Competitive Debate Arena, comprising 66 carefully selected Chinese debate motions. We recruit ten experienced human debaters and collect records of 200 debates involving Agent4Debate, baseline models, and humans. The evaluation employs the Debatrix automatic scoring system and professional human reviewers based on the established Debatrix-Elo and Human-Elo ranking. Experimental results indicate that the state-of-the-art Agent4Debate exhibits capabilities comparable to those of humans. Furthermore, ablation studies demonstrate the effectiveness of each component in the agent structure.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2024) studied this question.

synapsesocial.com/papers/68e5d23bb6db643587567e57https://doi.org/10.48550/arxiv.2408.04472
Ask AI
Helpful
Bookmark
Share
View Full Paper