Methodological audit uncovers eleven measurement failures during small language model benchmarking on consumer hardware, highlighting severe vulnerabilities in evaluation test suites.
Version 2 adds three cases recorded on the second day and splits the distribution table into five categories. Over roughly forty-eight hours of continuous measurement on a single consumer GPU, eleven claims were retracted. Each was recorded at the moment it died, in a fixed schema including a field that cannot be reconstructed afterwards: why it looked true. Five of the eleven were caused by the author's own instruments — three reporting events that had not happened, and two (new in this version) designed so that a correct answer scored zero. In one, distractors built as one-character perturbations of the gold identifier turned the distractor set into an error-correcting code for the answer: majority-voting them recovers it exactly, with no model, at k≥8. In the other, every line on the affected rung carried the same field marker, so the model's correct reading scored 0.00 against an incorrect key — a design that drives the score toward zero as reading improves. A third new case records an arm dropped from a published result without saying so: a fourth model whose 0.46 was 52% non-answers rather than interference sensitivity. Case identifiers (A-K) are stable and are not reassigned. The machine-readable ledger is included alongside the text.
No takes yet. Share an insight, caveat, or question.
satoshi akiyama (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: