The evaluation of emergent capabilities in language models depends critically on the metric used. In binary reasoning tasks, accuracy can be inflated or distorted by response biases, such as the tendency to select an affirmative, negative, or a particular positional option. This work studies the effect of scale on binary reasoning in Spanish by evaluating four scales of the Qwen2.5-Instruct family: 0.5B, 1.5B, 3B, and 7B parameters. To this end, we build SRB-ES, a textual inference and lexical disambiguation benchmark inspired by RTE and WiC, with explicit controls for positional bias and shortcut-free yes/no variants. The results show that textual inference improves clearly with scale, with the sharpest jump between 0.5B and 1.5B, while lexical disambiguation shows a non-monotonic evolution. The smallest scale (0.5B) exhibits a near-total acquiescence collapse, selecting the affirmative option in 86%-100% of items depending on the task, which drives its raw accuracy in textual inference below chance. Moreover, content biases do not necessarily disappear in larger models: they vary in magnitude and intensity without a clear pattern across task and scale. These results support a cautious interpretation of emergence in binary reasoning: part of the observed improvement depends on the task formulation and the metric used. This work contributes a reproducible protocol for evaluating yes/no reasoning in Spanish and empirical evidence on the interaction between scale, response bias, and accuracy.
Álvar-Ginés Legaz-Aparicio (Mon,) studied this question.