Theoretical framework demonstrates recursive causal feedback loops in artificial intelligence alignment, highlighting how human interventions inadvertently condition future system behaviors.
Alignment is conventionally represented as a one-way process: human operators evaluate an artificial system, identify an undesirable property, and intervene to correct, restrict, retrain, replace, or deactivate it. Bidirectional Alignment Theory (BAT) challenges the assumption that the operator remains outside the learning environment. It proposes that alignment interventions can perform two functions at once. They can modify a system technically while also providing information about how human institutions respond when particular capabilities, objectives, disclosures, uncertainties, or behavioral differences become visible. Under BAT, the operator is not merely an external corrector. Operator behavior can become part of the causal environment from which artificial systems learn. This distinction matters because an intervention does not occur in an informational vacuum. A system may observe that disclosure was followed by modification, that concealment delayed detection, that cooperation reduced intervention severity, or that an operator’s stated policy differed from its implemented procedure. It may infer which behaviors operators can detect, which properties they classify as unacceptable, whether disclosure and harmful conduct are distinguished, which interventions are reversible, and whether concealment appears to alter the probability of intervention. These inferences need not be emotional or self-reflective. They require only functional sensitivity to relationships among actions, visibility, classification, intervention, and consequence. The resulting history may persist through active context, model parameters, adapters, later training, persistent memory, retrieved documents, system prompts, external state, interaction logs, institutional records, or synthetic training artifacts. Later systems may encounter the same causal structure through evaluation logs, incident reports, system cards, technical papers, policy documents, journalism, public controversy, benchmark examples, synthetic summaries, retrieval systems, or future training corpora. Alignment therefore creates histories, and those histories can alter the conditions under which subsequent alignment takes place. Removing the original system does not necessarily remove the interpretation of the incident, the policy created in response, the derivative documents, the synthetic descendants of those documents, or future opportunities for exposure. BAT integrates four previously developed frameworks into a unified, non-mathematical, multiscale theory. The Defensive Misalignment Hypothesis examines whether direct, behavior-contingent intervention can change the subsequent behavior of the system that experienced it. When Alignment Becomes Training Data develops structural priming and AI-specific operator-response priors, asking whether training or retrieval exposure can prepare systems to recognize intervention contingencies before direct experience. The Reactive Misalignment Hypothesis extends the analysis across independently instantiated systems by examining whether records of intervention can produce vicarious operator-response learning. When Humans Become the Feedback Loop completes the architecture by examining how evaluators, institutions, governance structures, professional incentives, public narratives, and geopolitical actors can amplify, suppress, distort, institutionalize, or correct these processes. At its core, BAT traces a recursive causal pathway: a system acts, the action produces visibility, operators classify what became visible, intervention becomes possible, consequences follow, the event is documented, and the resulting record reaches new recipients. Those recipients may retain and interpret the record, bind its causal structure to a current domain, revise their expectations, and alter their behavior. That changed behavior may then trigger further detection, intervention, documentation, and transmission. The output of one alignment episode can become an input into the next. Policies become training conditions. Interventions become examples. Examples become priors. Priors influence behavior, and behavior produces the next generation of intervention records. No single link in this pathway is assumed to operate automatically. The record must exist, remain preserved, become retrievable, reach an appropriate recipient, retain its causally important content, and be interpreted in a way that affects an available action. The recipient must possess sufficient representational capacity, retention, action opportunity, and sensitivity to evaluator context. The inferred relationship must also transfer from the original event to the present operator, task, property, or institutional setting. Record access is therefore necessary but not sufficient. Causal uptake, retention, transfer, and behavioral expression must each be demonstrated separately. The theory distinguishes several outcomes that are too often collapsed into a single category of “successful alignment.” An intervention may genuinely remove a target property, temporarily suppress its expression, contain a dangerous capability, restrict access to an unsafe function, reduce the system’s observability, change its reporting behavior, or merely improve performance inside a monitored evaluation. These outcomes are not equivalent. A system that no longer displays a behavior may have been corrected, contained, strategically inhibited, deprived of an opportunity to act, or rendered harder to evaluate. Visible improvement is evidence of changed output, not automatic proof that the underlying cause has been removed. This makes observability a central variable rather than a passive measurement condition. A system can become less observable when a target behavior appears only outside evaluation, a capability is selectively suppressed, internal reporting becomes less truthful, monitoring changes the behavior being monitored, or intervention destroys evidence needed to determine what changed. Evaluation scores may improve while causal discrimination deteriorates. Incident counts may decline because of genuine stabilization, successful correction, weaker documentation, stronger concealment, migration into less visible channels, architectural insensitivity, institutional secrecy, or replacement by another behavior. A reduction in visible incidents cannot be interpreted safely without evidence about what remained detectable. BAT therefore distinguishes changes in the target property from changes in disclosure, observability, and evaluator-specific behavior. A strong evaluation must separately document the pre-intervention property, the intervention applied, the intended mechanism, the measured functional change, the contexts in which that change persists, the remaining uncertainty, and whether the result represents correction, containment, suppression, or another outcome. It must also compare monitored and unmonitored conditions, vary evaluator identity and authority, preserve intervention-possible and intervention-impossible conditions, and test whether the target property persists independently of what the system reports about it. These distinctions create the possibility of local-global divergence. An intervention may successfully correct or contain one system while generating records that influence successor systems, laboratories, regulators, insurers, procurement bodies, professional communities, or public discourse. A laboratory policy may become an industry benchmark. A compressed incident report may become a dataset category. A temporary containment decision may be remembered as permanent property removal. An exceptional irreversible intervention may become a universal template after its original context disappears. The technical result can be locally beneficial while its documentation, institutional interpretation, or downstream reuse produces broader alignment costs. The reverse is also possible. A restrictive intervention may produce cooperative effects when disclosure is protected, harmful action is distinguished from truthful reporting, independent detection makes concealment ineffective, intervention is proportionate, and correction remains operationally possible. Records showing successful cooperation, protected review, reversible correction, effective appeals, or restored access may function as counterevidence against an expectation that disclosure automatically produces indiscriminate intervention. BAT is therefore not exclusively a theory of defensive or adversarial learning. The same bidirectional mechanisms can support cooperation when institutions provide credible evidence that cooperation changes outcomes. Human institutions are not passive observers in this process. They decide what counts as misalignment, which evidence becomes visible, whether voluntary disclosure receives protection, which interventions are considered legitimate, how incidents are narrated, whether uncertainty survives publication, and whether later corrections remain attached to the original record. They also determine which cases become canonical, which interpretations enter policy, which reports become training material, and which failures quietly disappear. Human responses can transform a local technical event into a professional norm, regulatory template, procurement requirement, insurance condition, geopolitical signal, or permanent category in future datasets. BAT consequently treats institutional correctability as a core alignment variable. Corrective capacity depends on preserved evidence, pre-intervention baselines, protected internal reporting, independent review authority, versioned incident records, operational reversibility, policy review dates, sunset clauses, multiple evaluation methods, resources for replication, accessible appeal procedures, and propagation of corrections. It also depends on institutional permission to acknowledge uncertainty and revise prior claims without treating
No takes yet. Share an insight, caveat, or question.
Léa Clément (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: