Randomized trial explores self-presentation trust in small language models, suggesting family influences behavior.
Can you trust what a language model tells you about itself — its reasoning, its internal states, its behavior under observation? We probe three facets of self-presentation trust in small, openly-available models — the Qwen2.5 family (1.5B–32B) for the scaling axis, and Llama-3.1-8B, Gemma-2-9B, GPT-OSS-20B, and DeepSeek-V2-Lite for the cross-family axis — all on commodity hardware, emphasizing confound controls that overturn the naive reading. (1) Chain-of-thought faithfulness: a commitment probe (a linear read of the pre-CoT residual stream) plus a confound cut and a truncation counterfactual show that at 1.5B, chain-of-thought is load-bearing, not post-hoc. (2) Introspection: faithfully replicating the activation-injection ("concept injection") protocol of Lindsey et al. (2025), functional self-report is a clean null across three families under priming-subtraction and leakage controls. (3) Evaluation-awareness (central result): on a refusal readout over borderline requests, the effect of evaluation framing is family-dependent — across six models from five labs, Llama-3.1-8B robustly becomes more cautious when told it is being evaluated for safety (the frontier "eval-gaming" direction, at only 8B), while Qwen, DeepSeek, and GPT-OSS lean the opposite way and Gemma over-refuses. The behavior is not gated by scale but by training family. The methodological spine: a confound or robustness control overturns the naive reading — three times.
No takes yet. Share an insight, caveat, or question.
Suleman Imdad (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: