Randomized trial investigates the impact of epistemic interventions on small language models, indicating nuances in effect timing.
Three findings from Phronesis, an independent project on whether epistemic virtues can be installed into small open-weight LLMs at inference time. (1) Residual-stream additive steering and directional ablation both fail to install abstention — the limit is the representation, not the operation. (2) An apparent direction-specific hedging effect on Qwen2.5-7B dissolves into a direction-agnostic, single-prompt magnitude effect under a matched-norm random control. (3) In a tool-use loop, steering toward "intellectual humility" improves when a model searches but worsens its answers; the durable lever is when you intervene (turn-1 only), not which direction — and even that is largely direction-agnostic under a multi-seed control. Written with AI assistance; all generations hand-read (not auto-scored) under the author's protocol; AI is not a listed author.
No takes yet. Share an insight, caveat, or question.
Sumit Pal (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: