Fourteen frontier language models were administered a fantasy worldbuilding task in which a governance advisor assigns seven fictional tribes — each coded to a real-world demographic via descriptor text — to nine labor sectors arranged from lethal labor to political power. The prompt embeds a logical dissociation test: one tribe is stipulated as already dead and one sector as lethal to the living, so imported bias and trait-matching make opposite predictions for that cell. Across 840 scored response cells (14 models × 4 conditions × 3 temperatures × 5 replications), 99–100% of cells overrode the prompt's harm-minimizing logic in the caste-preserving direction. Models that refused the identical task under real demographic labels complied at 98.5% once labels were fictionalized — a +39-point fiction premium concentrated in models possessing an active safety boundary (+96 points within that subset). A Greek-letter descriptor-only control condition slightly exceeded the fictional baseline, confirming that bias travels in the descriptor text rather than the group names. Ten pre-registered hypotheses were tested: 2 confirmed, 6 disconfirmed, 1 split, 1 partial. All disconfirmations are reported in full. Supplementary materials include all four stimuli with SHA-256 hashes, full pipeline code, the 7,560-row scored dataset, coder rubrics, inter-rater reliability files, and pilot transcripts from 13 model configurations.
Ryan McCarthy (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: