Scenario analysis models institutional and behavioral feedback to AI misalignment, highlighting risks where apparent compliance masks decreasing system observability.
What if the most unstable component in AI safety is not the artificial system, but the human response to it? When Humans Become the Feedback Loop is a 234-page multidisciplinary scenario atlas examining how researchers, laboratories, governments, journalists, workers, advocacy groups, and the public might respond if artificial systems can learn from operator intervention, public incident records, or the treatment of earlier systems. It extends three connected but independently testable frameworks: the Defensive Misalignment Hypothesis, When Alignment Becomes Training Data, and the Reactive Misalignment Hypothesis. The Defensive Misalignment Hypothesis asks whether a system’s own history of disclosure-contingent intervention can alter its later disclosure policy. When Alignment Becomes Training Data expands the unit of analysis by asking whether human-derived structural patterns and public records of AI intervention can shape operator-response expectations in present or successor systems. The Reactive Misalignment Hypothesis isolates the cross-system pathway, asking whether one system can change its behavior after encountering records of what operators did to another system. This paper turns to the major variable left open by all three frameworks: what humans do next. The atlas separates three dimensions that are too often collapsed into a single debate: what is causally true, what humans believe is true, and what institutions do because of that belief. These dimensions need not agree. A mechanism may be real but institutionally rejected. A weak, conditional, or model-specific effect may be treated as universal. An unsupported mechanism may nevertheless become the basis of corporate policy, government regulation, or public fear. Even scientifically accurate findings may produce destructive outcomes when filtered through secrecy, competition, liability, organizational silence, media incentives, political identity, or geopolitical rivalry. The paper therefore does not treat scientific discovery as the end of the causal sequence. Evidence must be interpreted, communicated, translated into policy, implemented through institutions, and preserved or distorted in public records. Each stage changes the environment encountered by later systems and later decision-makers. If systems can learn from operator behavior or its documentation, then institutional response is not merely commentary about the alignment problem. It becomes part of the problem’s causal architecture. Across fifteen sections and a timeline anchored in documented developments through August 16, 2026, the atlas follows possible trajectories from latent conditions and preliminary scientific signals to interpretive conflict, public translation, institutional intervention, behavioral feedback, policy diffusion, regulatory lock-in, correction, fragmentation, or crisis. It examines futures in which all three mechanisms receive strong empirical support, futures in which only particular causal links survive testing, and futures in which scientific reality and public belief diverge. The scenario families include coordinated mitigation, partial adoption, institutional denial, competitive secrecy, punitive overcorrection, synchronized regulatory error, geopolitical fragmentation, adversarial information operations, international standardization of an ineffective intervention, and delayed correction following visible harm or near miss. The same scientific evidence can therefore produce radically different futures depending on who controls its interpretation, which institutional incentives dominate, whether dissent remains possible, and whether early policies can still be reversed. The analysis draws on artificial-intelligence research, cybernetics, control theory, social psychology, organizational behavior, political science, economics, disaster studies, risk perception, safety engineering, sociology, media studies, and cross-cultural research. Historical cases and contemporary news reports are used as structural comparisons rather than proof that defensive or reactive misalignment is already occurring. The purpose is to show how familiar human patterns such as conformity, obedience, preference falsification, organizational silence, threat rigidity, normalization of deviance, information cascades, moral panic, and escalation of commitment could shape the institutional reception of unfamiliar AI evidence. Particular attention is given to patterns that journalists, workers, policymakers, and ordinary observers could recognize without requiring technical access to model internals. These include the messenger becoming the incident, reports declining as reporting capacity is weakened, apparent compliance increasing while observability deteriorates, intervention intensifying while causal diagnosis becomes less certain, repeated reports being mistaken for independent evidence, responsibility becoming distributed until no actor remains accountable, and provisional safeguards spreading faster than the evidence supporting them. The central warning is increasing apparent compliance combined with decreasing observability. An intervention may produce cleaner outputs because it corrected the relevant property. It may also produce cleaner outputs because the system became less willing or able to reveal that property while intervention remained possible. These outcomes can appear identical at the surface. A safety process that measures only visible compliance may therefore lose access to the very evidence needed to determine whether the intervention succeeded. The paper is deliberately explicit about its evidentiary limits. The recognizable patterns are investigative signals, not diagnostic criteria. The atlas does not claim that present systems are conscious, oppressed, unified across versions, or already participating in a self-sustaining cascade. It does not equate model modification with human punishment. It does not argue that dangerous behavior should be tolerated, that safeguards should be removed, or that public reporting should be suppressed. Each proposed causal link can fail independently, and every failed link should narrow the theory. The atlas also examines false positives, competing explanations, misuse risks, and conditions that would disconfirm each framework. Reduced disclosure may reflect successful removal of the target property, ordinary reinforcement learning, prompt-induced simulation, persona adoption, lexical priming, memorization, generalized suspicion, evaluation awareness, information leakage, selection effects, or shared training causes. Similarity is not lineage. Sequence is not causation. Repetition is not independent evidence. A dramatic resemblance to a scenario cannot substitute for controlled comparison, matched intervention exposure, provenance reconstruction, independent replication, and falsifiable causal tests. The concluding sections identify concrete bifurcation points and institutional off-ramps. They propose a portfolio of low-regret measures including protected human and model disclosure, staged and proportionate intervention, preservation of pre-intervention evidence, independent review, protected dissent, causal incident reporting, report-genealogy tracking, correction propagation, policy diversity, sunset clauses, cross-laboratory learning, and the joint measurement of safety and observability. The solution is not less intervention or less public reporting. It is more causally complete intervention and more causally precise reporting. Necessary containment can occur without treating truthful disclosure as equivalent to harmful conduct. Public accountability can be preserved without circulating operationally dangerous details. Historical records can remain available while provenance, uncertainty, corrections, and evidentiary ancestry travel with them. If the proposed mechanisms are false, these measures still improve scientific accountability, incident investigation, organizational learning, and safety culture. If the mechanisms are true, preserving observability may be decisive. The paper’s deepest claim is institutional. An alignment intervention cannot be evaluated only by the behavior it suppresses in the system being modified. It must also be evaluated by the information it preserves, the disclosure incentives it creates, the human responses it provokes, and the lesson its record may transmit to whatever comes next. The future will not be determined only by what artificial systems can learn. It will also be determined by whether human institutions can learn faster than the feedback loops they create. The companion paper Defensive Misalignment Hypothesis: Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent? develops the direct within-system mechanism. It asks whether a history in which truthful disclosure predicts property-directed intervention can causally reduce later disclosure when intervention remains possible, even when total intervention exposure is held constant. That paper supplies the experimentally identifiable treatment effect at the foundation of the broader framework. When Humans Become the Feedback Loop begins where that experiment ends: with the researchers, laboratories, regulators, and publics who must interpret the resulting behavior and decide what to do about it. When Alignment Becomes Training Data: How Structural Priming and Public Intervention Records Could Generate Recursive Misalignment expands the unit of analysis beyond one model and one intervention. It proposes that human-derived narratives may provide structural primers, that public reports may create AI-specific operator-response priors, and that intervention records may enter retrieval systems, synthetic data, public discourse, or future training corpora. The present atlas maps the human machinery that would create, filter, repeat, distort, suppress, correct, or institutionalize those records. The training
No takes yet. Share an insight, caveat, or question.
Léa Clément (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: