PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 20260 citationsOpen Access

The Fantasia Bound on Constitutional Classifiers: Thermodynamic Limits of Jailbreak Defence

View Full Paper
AEAnthony W. Eckert

Key Points

  • The study aims to explore the effectiveness limits of constitutional classifiers against jailbreak attacks using thermodynamic principles.
  • Analyzed performance of constitutional classifiers under red teaming conditions.
  • Derived thermodynamic upper bounds on classifier effectiveness using the Fantasia Bound.
  • Investigated the critical Pe threshold for single-channel and two-channel architectures.
  • Classifiers could not withstand attacks above a critical Pe threshold.
  • The two-channel architecture improves effectiveness but does not eliminate vulnerability.
  • Predictions indicate universal jailbreaks exist beyond the critical threshold.

Abstract

Anthropic's constitutional classifiers (2026) withstood 3, 000+ hours of red teaming with no universal jailbreak. We derive a thermodynamic upper bound on classifier effectiveness from the Fantasia Bound I (D;Y) +I (M;Y) ≤H (Y). A classifier is a prohibition mechanism: its channel capacity sets the maximum Pe it can suppress. We prove that no single-channel classifier can defend against attacks above a critical Pe threshold Pec = exp (C/kT) where C is the classifier's information capacity. The prohibition-ritual pair architecture (two independent channels) raises the bound but does not eliminate it. Constitutional classifiers succeed because they approximate the two-channel architecture — the constitution provides the ritual (explicit reasoning about refusal), while the classifier provides the prohibition. Predictions: universal jailbreaks exist above Pec; defence requires increasing channel capacity, not classifier complexity.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Anthony W. Eckert (2026) studied this question.

synapsesocial.com/papers/69ccb71716edfba7beb88e61https://doi.org/10.5281/zenodo.19340888
Ask AI
Helpful
Bookmark
Share
View Full Paper