Theoretical analysis and preference data audit reveals structural non-identifiability in reward models, highlighting deployable criteria to prevent reward hacking in artificial intelligence.
Uncertainty about reward functions, world models, and agent capabilities in AIalignment is usually treated as parametric: given enough data, posteriors are assumedto concentrate. We show that much of this uncertainty is structural—thelikelihood is invariant under a nontrivial transformation group, so the data determineonly an identified set. The identifiability facts we use are classical (affinenon-identifiability of Bradley–Terry utilities; the connectedness criterion for theiridentification; orbit non-identifiability of group-invariant likelihoods). Our centralcontribution is a manipulation-cost criterion that separates safety constraints robustto observational manipulation (invariant constraints, positive cost) from gameableones (non-invariant, zero cost), together with a deploy-time check obtained byreinterpreting the classical comparison-graph connectivity condition as a rewardhackingvulnerability criterion. We also show that endogenous world-model nonidentifiability,reward hacking, and the undetectability of mesa-objectives are instancesof a single partial-identification structure. Five reproducible experimentsvalidate the finite-sample behavior of the framework; an item-level audit of a publicRLHF preference dataset confirms, as structurally expected, that tabular preferencedata leave essentially all cross-response comparisons unidentified. We positionthe work as a synthesizing framework with a practical safety criterion, not as newidentifiability theorems.
No takes yet. Share an insight, caveat, or question.
Yavuz Selim Kılınç (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: