PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 11, 20260 citationsOpen Access

Safety-Constrained Architecture Development: Preregistered Causal Experiments and Bounded Scientific Ledgers for AI Systems

View Full Paper
VZVlad Zotta

Key Points

  • This research aims to develop a framework for safety-constrained AI architecture to improve auditability and design integrity.
  • Developed a repository-native framework with six core mechanisms for safety oversight.
  • Conducted three experiments to apply the framework in a private program assessing AI systems.
  • Implemented a family-termination rule and other measures to enhance future program governance.
  • Improvements in a matched-readout intervention were made but exceeded safety limits, leading to rejection for operational use.
  • Subsequent interventions identified specific safety issues while being rejected as operational based on performance.
  • An independent evaluation identified additional weaknesses at a program level, prompting governance improvements.

Abstract

Exploratory AI architecture research gives the same team unusually broad control over benchmark construction, implementation, metric selection, threshold choice, reruns, and interpretation. That concentration of discretion creates opportunities for outcome-driven design changes, selective reporting, test-set leakage, overclaiming, and thedisappearance of failed directions. We present a repository-native framework for safety-constrained architecture development intended to make those failure modes visible and costly before external peer review is available. The implemented framework combines six core mechanisms: preregistration-first commit chronology with auditable ancestry; blind-evaluation gates with explicit unlock criteria and access counters; adjacent-contrast causal ladders that change one component at a time; bounded scientific ledgers that store claims together with their tested limits; a typed failure protocol that preserves negative results and resurrection conditions; and canonical numerical identities paired with executable decision invariants.We illustrate those six mechanisms with a three-experiment development-only sequence from Program K, a private program studying consequence-aware epistemic mechanisms. A matched-readout intervention improved several utility measures but breached a preregistered critical-safety allowance and was not accepted for operational use.Subsequent single-variable interventions localized the safety failure first to a sufficiency channel and then to cross-field applicability within that channel, while again rejecting the diagnostic intervention as an operational architecture because it created excessive false gaps. The sequence also exposed a provenance ambiguity in themiddle experiment: its scientific rules were fixed in an authoritative repository issue before evaluation, but its branch history did not preserve a separately auditable preregistration-only commit. That deficiency was recorded and led to a stricter chronology rule that the final experiment then satisfied. After the case sequence, an independent read-only evaluation identified additional program-level weaknesses. On 31 July 2026 the roadmap prospectively adopted a family-termination rule, a strong contemporary-baseline checkpoint, a preregistered naturalistic-data probe, external adversarial review, and explicit safety-coverage and cost reporting. These amendments did not govern the historical case study and are presented as review-driven hardening, not retroactive evidence of prior practice. We present the overall record as evidence of auditability, bounded causal attribution, and process correction, not as evidence that the governance framework improves scientific outcomes or that Program K generalizes beyond its synthetic benchmark.This record contains the preprint manuscript. The accompanying sanitized evidence bundle is available at https://doi.org/10.5281/zenodo.21726365.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vlad Zotta (2026) studied this question.

synapsesocial.com/papers/6a7ace273401087f2249de91https://doi.org/10.5281/zenodo.21861424
Ask AI
Helpful
Bookmark
Share
View Full Paper