PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 27, 20250 citationsOpen Access

The AI Alignment trilemma: Inner Alignment, Outer Alignment and Capability Cannot be Simultaneously Maximised

View Full Paper
RJReich, Jonathan

Key Points

Key points are not available for this paper at this time.

Abstract

The AI Alignment Trilemma: A Formal Impossibility Result This paper proves that AI systems face a fundamental trilemma: inner alignment (mesa-optimizer control), outer alignment (reward specification correctness), and capability (model power) cannot be simultaneously maximized. We establish the constraint I² + O² + C² ≤ 1, demonstrating that "aligned superintelligence" is mathematically impossible. The result is derived from finite optimization budgets and empirically measured quadratic scaling of capability costs, analogous to the CAP theorem in distributed systems. All proofs are machine-verified in Lean 4 (738 lines, 0 sorrys), and empirical predictions are validated by convergent evidence from all three major frontier AI labs: OpenAI o1 (alignment faking, reward hacking), Anthropic Claude Opus 4.5 (specification gaming), and Google Gemini 3 (manipulation propensity scaling with capability). This shifts the AI safety research paradigm from "solving alignment" to navigating the Pareto frontier of achievable tradeoffs, with direct implications for safety policy and capability regulation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Reich, Jonathan (2025) studied this question.

synapsesocial.com/papers/694031da2d562116f2907229https://doi.org/10.5281/zenodo.17739090
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The AI Alignment trilemma: Inner Alignment, Outer Alignment and Capability Cannot be Simultaneously Maximised2025
  2. 2There and Back Again: The AI Alignment Paradox2024 · 1 citations
  3. 3AI alignment boundaries2025
  4. 4The Architecture of Acceptable Consequences: A Constraint-Based Proposed Solution to the AI Alignment Problem2026
  5. 5Safety as Natural Emergence: The Conciseness Framework as a Foundation for Intrinsic AI Alignment2026