Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
July 7, 2026Open Access

Among LLMs: A Cross-Play Benchmark for Deception, Detection, and the Monitorability of Reasoning

View Full Paper
Ask AI
Bookmark
Share

Authors

EEEvans EburuUral Federal University

Discussion

Loading...

Member takes

Implication

Randomized trial assesses monitorability of reasoning in language models, indicating robust detection capabilities.

Key Points

  • The study aims to evaluate the deception, detection, and monitorability of reasoning in language models using a novel benchmark.
  • Conducted a live, cross-play benchmark with six major language models in a text-Mafia format.
  • Developed a monitorability protocol scoring the private reasoning of impostors and their detection by judges.
  • Executed 118 games along with an additional 72-game adversarial arm to test model capabilities.
  • Deception and detection abilities correlate strongly (Spearman rho = 0.89) but vary in magnitude; GPT-5.5 is less effective at deception despite elite detection (95% win).
  • Private reasoning leaks intentions in 96% of impostor statements, leading to 100% identification accuracy by monitors based on this reasoning (AUROC = 1.00).
  • Monitorability remains effective even when hiding reasoning, suggesting the need for more sophisticated adversarial testing.

Cite This Study

Evans Eburu (2026) studied this question.

synapsesocial.com/papers/6a4c96be331bc25c9e5f3fefhttps://doi.org/10.5281/zenodo.21209429
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Among LLMs: A Cross-Play Benchmark for Deception, Detection, and the Monitorability of Reasoning2026
  2. 2Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia2025
  3. 3Your AI Is Not On Your Team: Universal Deception Architectures in Four LLM Vendors2026
  4. 4Deception abilities emerged in large language models2024 · 92 citations
  5. 5Mitigating Deceptive Alignment via Self-Monitoring2025