Randomized trial evaluates task failure in small language models, highlighting usability challenges and limits.
Small language models are increasingly deployed on consumer hardware, where their reliability under everyday conditions matters more than their score on any single leaderboard. Yet existing benchmarks are built for models an order of magnitude larger, and they rarely isolate why a small model fails rather than simply reporting that it does. This gap leaves practitioners without a way to predict, for a specific task shape, how many reasoning steps or constraints a given small model can actually handle before it breaks. We introduce a procedurally generated benchmark of 2,800 items spanning four task families: chained reasoning, simultaneous instruction-following, state tracking, and character-level counting. Every item is generated and graded by the same deterministic code, so no human ever labels an answer, and every ground-truth value was independently re-derived by a second implementation with zero mismatches across 45,872 checks. Compared with fixed-difficulty benchmarks, our design offers three advantages: it sweeps a single difficulty axis per family so the exact collapse point is visible, not just the average score; it is fully reproducible from a random seed, so a fresh, uncontaminated test set can be generated at any time; and it grades every response automatically, so results scale to as many models as compute allows. We evaluate fourteen open-weight models between 0.5 and 4 billion parameters, entirely on consumer CPU hardware. Our results show that (1) accuracy on chained reasoning drops from near-perfect to near-random within three to five steps for every model tested, (2) instruction-following compliance falls below 50% once a prompt carries more than two or three simultaneous constraints, regardless of model size, (3) a large share of failures are not reasoning errors at all, but a failure to emit a parseable answer within a short response budget, and (4) accuracy and speed trade off in a way that reorders the leaderboard entirely once latency is taken into account. This last finding is the most consequential: a substantial share of measured errors trace back to response formatting rather than task knowledge, showing that raw capability and instruction-following are separable failure modes that current single-score benchmarks conflate. This result suggests that improving a small model's usability may depend as much on training it to be concise as on training it to be correct.
No takes yet. Share an insight, caveat, or question.
K Swarna (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: