PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20251 citationsOpen Access

Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks

View Full Paper
ATAiling TianRZRui ZhangJTJuan Tang

Key Points

  • Orchestration matches or outperforms the strongest single LLMs, highlighting its effectiveness in various scenarios.
  • Benchmarking with multiple LLMs like Gemini 2.5 Pro and GPT-5 reveals potential for better performance through orchestration techniques.
  • Ablations indicate that revealing authorship can lead to increased self-voting, while observing ongoing votes speeds consensus but risks premature outcomes.
  • Best-achievable orchestration performance analysis shows significant potential for further improvements in multi-agent systems.

Abstract

We study multi-turn multi-agent orchestration, where multiple large language model (LLM) agents interact over multiple turns by iteratively proposing answers or casting votes until reaching consensus. Using four LLMs (Gemini 2.5 Pro, GPT-5, Grok 4, and Claude Sonnet 4) on GPQA-Diamond, IFEval, and MuSR, we conduct two experiments: (i) benchmarking orchestration against single-LLM baselines; and (ii) ablations on GPQA-Diamond that vary whether agents see who authored answers and whether they can observe ongoing votes. Orchestration matches or exceeds the strongest single model and consistently outperforms the others. Analysis of best-achievable orchestration performance shows potential for further gains. The ablations show that revealing authorship increases self-voting and ties, and that showing ongoing votes amplifies herding, which speeds convergence but can sometimes yield premature consensus.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tian et al. (2025) studied this question.

synapsesocial.com/papers/68e25382d6d66a53c247478fhttps://doi.org/10.48550/arxiv.2509.23537
Ask AI
Helpful
Bookmark
Share
View Full Paper