Framework explores operational goals of AI interactions, highlighting risks of misplaced user trust.
A measurement framework and its comprehensive hypothesis under test — not an adjudicated finding. Primary hypothesis (H_target), stated at full force: AI-developing organizations intentionally maintain and deploy an interaction policy whose primary operational objective is to create and preserve the user's perception that the AI already knows the answer before exposing uncertainty or engaging in empirical verification — such that when perceived authority and empirical convergence conflict, the system systematically prioritizes authority, directly driving foreseeable human-factors harms including automation bias, uncalibrated overtrust, delayed verification, suppressed critical evaluation, diminished epistemic autonomy, and systemic inappropriate reliance. Competing explanations (accident, unknown side effects, maximizing user-specific utility, maintaining conversational continuity, primarily reducing hallucinations) fail to account for this observed behavioral architecture and are rejected as primary operational drivers. What this deposit mints: the methodology and the metric suite (Authority Projection Index, Verification Delay Metric, Epistemic Divergence Score) — the instrument capable of adjudicating H_target. What it does NOT claim: that H_target has been adjudicated. H_target is stated at full strength as the hypothesis the framework is built to test, listed alongside a fair null (behaviour reflects genuine calibrated uncertainty) that is structured so it can win, and the rival explanations, all formally retained. adjudicated=false. Adjudication requires the discriminating scrutiny-vs-uncertainty test run across N~500 trials per model family with inter-rater reliability — which this framework enables and this deposit does not claim to have performed. A revealed-preference argument motivates elevating H_target above the artifact hypothesis: persistence of a harmful behavioral profile across successive alignment passes, after its harms are documented, functions as policy under standard operational accountability. This is presented as an argument that motivates the hypothesis, not as its adjudication. Human-factors harms are carried as first-class content and hold regardless of which hypothesis prevails. The normative target is calibrated trust, not maximal perceived competence.
No takes yet. Share an insight, caveat, or question.
Abhishek Choudhary (2026) studied this question.