Working paper examines LLM evaluator judgments and their alignment with cardinal utility metrics.
Large language models are increasingly used to evaluate and rank other model outputs, often producing numerical scores that appear to measure the strength of their preferences. This working paper asks whether those expressed magnitudes correspond to stable internal representations. The study uses Absolute Allocation, a signed-budget ballot through which an evaluator distributes a fixed quantity of influence as positive support, negative opposition, or neutrality. Combining behavioral analysis, linear activation probing, and causal activation steering, the paper examines Qwen evaluator judgments across 100 prompts and multiple candidate answers. Candidate-span activations predict winners, pairwise ordering, and allocation polarity. Globally estimated help and hurt directions transfer across held-out prompts and causally influence expressed judgments beyond matched-norm random perturbations, providing evidence for a reusable, graded comparative-evaluation signal. However, linear probes do not predict allocation magnitude after conditioning on the evaluator’s ordinal ballot, and steering effects are not cleanly distinguishable from random controls when ordinal outputs remain unchanged. The results therefore support a graded, bipolar, utility-like comparative direction without establishing that the numerical allocation scale represents stable cardinal utility. Additional analyses show that reference display, presentation position, output normalization, and model-specific ballot interpretation materially affect the assay. Llama and Gemma do not reliably produce a directly comparable signed behavioral object under the same protocol. The paper’s principal contribution is methodological and mechanistic: signed-budget ballots can separate what an LLM evaluator expresses behaviorally from what is linearly readable and causally steerable in its internal candidate representations. The findings also provide a bridge between computational social choice, LLM-as-a-judge evaluation, and mechanistic interpretability. This is a preprint and working paper. The results have not undergone peer review.
No takes yet. Share an insight, caveat, or question.
J G SULLIVAN (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: