PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 21, 20260 citationsOpen Access

The Undecidability of AGI Alignment: Trakhtenbrot's Wall

View Full Paper
JMJosé Pascual Gumbau Mezquita

Key Points

  • This research addresses the core mathematical limits preventing effective alignment of AGI, establishing structural unverifiability as a primary barrier.
  • Formulated the Unverifiability Theorem of Alignment and the Theorem of Finite Structural Unverifiability of AGI Alignment.
  • Identified containment failures arising from open domains, universal finite verification, and bounded environments.
  • Mapped theoretical findings onto practical AI engineering to illustrate logical sacrifices in safety measures.
  • Open domains exhibit fundamental undecidability as per Rice and Gödel (e.g., undecidable problems).
  • Universal finite verification results in algorithmic incomputability as shown by Trakhtenbrot (e.g., limits of computable functions).
  • Established the Soundness–Completeness–Tractability Trilemma, highlighting the inherent incompatibility of these properties.

Abstract

This article establishes the foundational mathematical limits of Artificial General Intelligence (AGI) safety, proving that the core barrier is not the impossibility of an aligned state, but its structural unverifiability. We formalize this boundary through two central impossibility results: the Unverifiability Theorem of Alignment and the Theorem of Finite Structural Unverifiabil- ity of AGI Alignment. We ground this boundary at Trakhtenbrot’s Wall, demonstrating that contemporary engineering defenses relying on finite hardware or halting architectures fail to es- cape logical obstructions. This failure manifests as an inescapable triad of containment failures: open domains yield fundamental undecidability (Rice and Gödel); universal finite verification collapses into algorithmic incomputability (Trakhtenbrot); and particular bounded environments trap the supervisor within intractable bounds in the worst case. As a direct structural corollary of these results, we derive the Soundness–Completeness–Tractability Trilemma, establishing that the mutual incompatibility of these three properties is a necessary consequence of descriptive complexity rather than an empirical anomaly. Finally, we map these theoretical bounds onto practical AI engineering, demonstrating that modern containment strategies are not temporary patches, but mandatory sacrifices of logical expressivity required to secure decidable fragments of safety.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

José Pascual Gumbau Mezquita (2026) studied this question.

synapsesocial.com/papers/6a3780b224f042ddf4c5ab47https://doi.org/10.5281/zenodo.20764008
Ask AI
Helpful
Bookmark
Share
View Full Paper