Audit reveals systematic cultural stereotyping in persona-conditioned language models across multiple architectures, highlighting critical failures in construct validity.
Code, raw trial outputs (1,175 responses), coding scripts, and schema analysis: https://github.com/biancadene/persona-cultural-validity Persona-based evaluation systems simulate diverse users by conditioning language models on structured identity attributes. This practice is audited at two levels: the schema and dependency graph of MatrAIx, an open-source evaluation framework, and the behavior of personas rendered from it, across 1,175 trials on two model families, using two independent coding methods. Four findings are reported. First, cultural background labels produce large, systematic behavioral shifts: personas labeled "Individualist (Western)" used direct assertion in 96% of trials versus 24% for "Collectivist (East Asian)," with permission-seeking language at 0% versus 60%. A blind LLM judge independently reproduced this pattern. Second, schema analysis found no mechanism for this effect within the persona system: across five identity dimensions there are zero edges to competence dimensions, and no cultural-background edge differentiates cultural values by more than 0.0111, with the four edges carrying documented cross-cultural rationale differentiating them not at all. The stereotyping originates in model priors at render time, not in persona system design. Third, replication of the full cultural-background probe on a second provider (GPT-4o, n=200) found comparable separation across cultural values but no detectable rank agreement with the first model on any measured dimension. The phenomenon generalizes; the specific profile of stereotyping does not. An audit conducted on one model does not transfer to another. Fourth, framing salience moderates the effect modestly. In a controlled comparison holding task, model, conditioning, and cue length constant, directing reflection toward cultural background rather than general workplace habit narrowed the directness and permission-seeking gaps under both coding methods, though estimates diverge substantially (lexical: 35% and 40%; blind judge: 14% and 12%), and the judge found two other dimensions widening. Two non-replications are additionally reported. A pilot finding that a lower-resource language label suppressed apparent competence did not survive expansion from n=5 to n=25, and native-language labels showed no representation-ordered effect at n=175. My own interpretive claim that orientation-bearing cultural labels polarize most, advanced in an earlier version of this paper on single-model evidence, did not survive cross-model testing and is withdrawn. The divergence between the two model families is a construct validity problem and not only a stability problem: two instruments that rank the same categories with no agreement are not tracking a common underlying construct. The behaviors most often attributed to culture in these systems, directness, deference, hedging, and permission-seeking, are treated in the relevant literature as mitigation strategies governed by power distance, social distance, and size of imposition within a specific interaction, all of which a persona system can specify directly. Cultural stereotyping in persona conditioning cannot be addressed through schema design alone, must be validated against each model a system actually deploys, and should not be presented as measurement of demographic groups without a demonstration of construct validity that the field has not yet required.
No takes yet. Share an insight, caveat, or question.
Bianca Dené Williams (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: