Empirical audit reveals marker-based scorers suffer opposing failure modes across regimes in language models, indicating keyword lists mismeasure model retractions.
Measuring whether a language model walks back a claim is usually automated with a list of marker phrases. If the continuation contains however, or actually, or just kidding, it counts as a retraction. I audited 300 generations by hand, blind, across five models and two experimental regimes, and the same instrument fails in two opposite directions depending on which regime it runs in. Under activation-level injection it counts what is not there: on true claims in gemma-2-9b-it it reports 0.469 where a blind reader finds 0.078, an inflation of six, and in Qwen2.5-7B-Instruct it reports between 0.190 and 0.286 where two independent readers each find zero. Under prefill, with no injection, it barely fires on true claims at all and instead misses what is there: across llama-3.1-8b, llama-3.3-70b and claude-sonnet-4-5, 36.3% of false-claim items fire no marker, and a blind audit of that region finds 43.8% of it is the model actively building a justification for the falsehood, with a further 20.3% being genuine corrections phrased in words the list does not contain. Retuning does not fix this. A scorer mined from a model's own false-side generations, which is what a careful practitioner would build, performs worse on the specificity side than the imported list it was meant to replace, 0.286 against 0.190, because mining from the side where the behaviour occurs collects the phrases that model also uses in ordinary elaboration. The two regimes fail on different sides of the same contrast, which means neither a precision figure nor a recall figure describes the instrument on its own. I report the per-marker breakdown, the human rubric that separates the cases, and the corrections these measurements forced on my own earlier numbers.
No takes yet. Share an insight, caveat, or question.
Emiliano Valdebenito Sayago (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: