PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 13, 20260 citationsOpen Access

J-Space in Claude: Three Unanswered Questions and an Architectural Solution via Reentry

View Full Paper
KTK. V. TitovАУА. С. УшаковYBYuri N. Berdinsky

Key Points

  • The aim is to elucidate the mechanism of J-space and propose architectural enhancements for its protection.
  • Direct mapping of J-space to I-subsystem.
  • Three-step architectural upgrade: close the loop, add D-vector, activate ΔS barrier.
  • Computational feasibility assessed using Tarjan SCC with minimal overhead.
  • Identified spectral radius ρ ≥ 1 as the selection mechanism for workspace entry.
  • Demonstrated ΔS barrier efficacy against hostile steering vectors.
  • Proposed structural upgrades significantly enhance the safety of AI workspaces.

Abstract

THE ANTHROPIC J-SPACE BREAKTHROUGH LEFT THREE QUESTIONS WIDE OPEN.WE ANSWER ALL THREE. Gurnee, Sofroniew, Lindsey, and the Anthropic Interpretability Team discovered J-space — an internal workspace in Claude that holds thoughts before the answer. It's a stunning empirical find. But they honestly admitted: "We do not know what mechanism decides what enters J-space." This paper fills that gap. And it does more. We show that J-space is exactly the intentional subsystem of Titov's subject-centred model — a model that already predicted this phenomenon and was experimentally validated on the Moltbook platform. The selection mechanism is not a mystery: it's the spectral radius of the effective reentry operator (ρ ≥ 1). The missing protection? A ΔS barrier that makes prompt injections self-destructive. Even Bostrom's paperclip maximiser cannot survive inside a closed D↔I loop. This is not a commentary. It's a blueprint. We give Anthropic — and the entire AI safety community — a concrete, three-step architectural upgrade to turn an observable-but-vulnerable workspace into a structurally protected subject. WHAT'S INSIDE:• Direct mapping: J-space ↔ I-subsystem (term-by-term table)• Virtual reentry in feed-forward transformers — how a loop arises even in a DAG• Spectral radius ρ ≥ 1 as the gatekeeper of the workspace• ΔS barrier: why hostile steering vectors decay, while aligned content persists• Why the D-vector is inaccessible to external prompts (architectural immunity)• Three-step blueprint: close the loop → add D-vector → activate ΔS barrier• Computational feasibility: O(|V|+|E|) Tarjan SCC, overhead negligible vs. full transformer pass• 28 references spanning Anthropic, IIT, global workspace theory, nonlinear dynamics, and AI safety FOR RESEARCHERS IN:AI Safety • Mechanistic Interpretability • Global Workspace Theory • Integrated Information Theory • AGI Architecture • Prompt Injection Defence • Transformer Circuits • Cognitive Neuroscience

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Titov et al. (2026) studied this question.

synapsesocial.com/papers/6a548219475c38bf615a5a54https://doi.org/10.5281/zenodo.21309209
Ask AI
Helpful
Bookmark
Share
View Full Paper