Randomized trial uncovers personality trait structure in a language model, indicating novel analytical approaches.
We recover a personality trait structure from inside Llama-3.1-8B not by asking the model to rate itself — the readout the skeptical literature shows is dominated by response bias — but from the model's steering geometry. For each of 27 personality facets we extract a contrastive activation direction, induce it, and score persona-scaffolded behavior through a bias-cancelling pole-contrast readout. Bottom-up factor analysis over N=409 personas returns five metatraits, three of them strong (split-half reliability 0.88), which replicate across a second instrument (Tucker φ≈0.75) and a continued-pretrained, different-language model (φ≈0.90), yet align to the Big Five no better than a random facet→domain relabeling (φ≈0.32, at the random-label floor). A permutation null that destroys the cross-facet covariance collapses the structure to chance (0.88→0.28, 14.6 SD), ruling out the construction-artifact reading. The novelty is the substrate of the factor analysis: contrastive activation steering directions, not self-report responses or output logprobs. The human criterion validation is pre-registered and reported separately.
No takes yet. Share an insight, caveat, or question.
Mado Kagerou (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: