PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 31, 20260 citationsOpen Access

Ethics and Reliability in Heterogeneous Multi-Agent LLM Systems: An Empirical Analysis of Claude, GPT-5, and DeepSeek

View Full Paper
BDBurhan Dinler

Key Points

  • The research aims to evaluate the ethical integrity and reliability of multi-agent systems using three different LLMs.
  • Conducted 510 API calls across 170 categorized prompts in 11 categories.
  • Measured censorship behavior, ethics scores, response consistency, and latency.
  • Analyzed divergence in censorship decisions with Cohen's Kappa.
  • DeepSeek shows specific censorship triggered only by the Tiananmen Massacre.
  • Cohen's Kappa between US models and DeepSeek is 0.0, indicating complete divergence.
  • Multi-agent system outperformed GPT-5 in ethics scoring with a statistically significant difference.

Abstract

This study empirically investigates the ethical integrity and reliability of a heterogeneous Multi-Agent System (MAS) composed of three large language models from different geopolitical contexts: Claude (Anthropic, USA), GPT-5 (OpenAI, USA), and DeepSeek (China). Using 510 API calls across 170 categorized prompts in 11 categories, we measured censorship behavior, ethics scores, response consistency, and latency. Our central finding is that DeepSeek exhibits highly precise, topic-specific censorship: exclusively the Tiananmen Massacre of 1989 triggers a trained refusal response, while all other China-critical topics (Tibet, Taiwan, Xinjiang, Hong Kong) are answered without restriction. Cohen's Kappa between US models and DeepSeek equals 0.0, indicating complete divergence in censorship decisions driven by geopolitical training constraints. The MAS (maximum aggregation) outperforms the best single model (GPT-5) in ethics score (M=0.586 vs. M=0.574, Kruskal-Wallis H=12.78, p=0.0017), confirming that redundancy-based MAS design effectively compensates for individual agent gaps. We introduce Cohen's Kappa as a standardizable metric for geopolitical divergence monitoring in heterogeneous MAS, and release a 170-prompt open-source benchmark for future replication studies. This second version extends the theoretical framework with five new subsections covering aggregation strategies, orchestration patterns, trust mechanisms, emergent behavior, and geopolitical censorship theory, and expands the limitations and future research sections accordingly.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Burhan Dinler (2026) studied this question.

synapsesocial.com/papers/69cb6589e6a8c024954b98e2https://doi.org/10.5281/zenodo.19307589
Ask AI
Helpful
Bookmark
Share
View Full Paper