Randomized trial investigates the effects of source attribution on language models, revealing the hazards of false attribution.
Correct source attribution measurably improves a language model's prediction of the text that follows it, and the effect does not depend on the model having seen that source before. Measured against a proper baseline, however, the more consequential finding runs the other way: training on FALSE attribution destroys attribution ability the model already had, at roughly twice the magnitude that correct attribution adds it. Eleven GPT-2 small (124M) models were fine-tuned on an identical set of 50,000 Wikipedia sentences drawn from 2,589 articles. Three conditions differ only in the tag prepended to each sentence: the true article title, a false title drawn from the same pool, or no tag at all. Those conditions are crossed with two ways of partitioning the same corpus, and every unseen-source result is replicated across independent seeds. Hyperparameters and data ordering are held constant. The false-tag condition uses a derangement rather than a plain shuffle, so no sentence retains its own title and the distribution of titles is preserved exactly. EXPERIMENTS 1 AND 2 — CONDITIONING Against the false-tag control, which holds tag length, format, and title distribution fixed and varies only whether the tag is true, a correct tag is worth 0.225 nats/token when the source appeared in training and 0.223 nats/token when it did not. Under the article-level split, whole articles are assigned to one side or the other, so 100% of test titles are never seen during fine-tuning. The advantage did not move — a difference of 0.002 nats. On the sentence's first word, where the tag's influence is strongest, the advantage is larger on unseen sources (1.36 nats) than on seen ones (1.14). Provenance conditioning is therefore not a memorized lookup from title to content. The unseen-source result is replicated across three provenance seeds and three scrambled seeds (the latter with independent derangements). Arm means are 3.4318 +/- 0.0008 and 3.6551 +/- 0.0004 with disjoint ranges; the advantage varies only between 0.2221 and 0.2244 across all nine seed pairings, roughly 280 times the observed seed standard deviation. EXPERIMENT 3 — SOURCE RECOVERY, AND THE DAMAGE FROM FALSE ATTRIBUTION The prefix design trains p(text | source), but attribution proper is the inverse. By Bayes, p(source | text) is proportional to p(text | source) times p(source), so the models can be inverted into attributors with no retraining. Each of 2,000 held-out sentences was ranked against 49 distractor titles. The provenance models identify the true source of an unseen sentence 75.4% of the time, and 52.9% of the time even when the sentence shares no content word with its title — 26 times chance. The mechanism is topical, not lexical. That figure needs a baseline, and supplying one changes its meaning. A model trained with no tags at all recovers sources at 64.7%, because GPT-2's pretraining already links title strings to the content they head. Most of the apparent attribution ability is inherited, not taught. Against that floor: training on TRUE attribution: +10.7 points top-1 (+12.8 zero-overlap) training on FALSE attribution: -25.8 points top-1 (-22.1 zero-overlap) Corrupted provenance is 2.4 times as damaging as correct provenance is helpful on the full sample, and 1.7 times on the zero-overlap subset. The scrambled model did not merely fail to learn from bad metadata; having learned the tag channel was unreliable, it discarded that channel and lost pretrained capability with it. Both figures are replicated. The scrambled arm was retrained with fresh derangements and fresh training seeds (three seeds; the new wrong-title assignments overlap the original on under 0.04% of sentences), and the provenance and control arms with fresh training seeds. Per-arm seed ranges are disjoint: provenance 74.70-75.90, control 64.50-64.85, scrambled 37.85-39.40. Across all six scrambled-versus-control pairings the damage ranges from -25.1 to -27.0. Critically, none of this is visible in predictive loss. Training losses agree to within 0.0013 nats across seeds within each condition, and the corrupted model slightly outperforms the untagged control on masked sentence loss. Predictive quality and attribution integrity come apart, and a system evaluated on perplexity alone would not notice the damage. OPEN-SET ATTRIBUTION The ranking above is closed-set: the true source is always among the fifty candidates, so the model need only rank, never abstain. Measuring the harder question — is the true source present at all? — each sentence was scored against two equal-sized candidate sets, one containing the true source and one not, discriminated by a scale-free statistic z = (max - mean) / sd over the pool. provenance AUC 0.7846 +/- 0.0016 rejection acc 72.96% zero-overlap AUC 0.6324 control AUC 0.7398 +/- 0.0007 rejection acc 69.59% zero-overlap AUC 0.6038 scrambled AUC 0.6387 +/- 0.0017 rejection acc 60.81% zero-overlap AUC 0.5507 Two qualifications follow. Open-set attribution is substantially harder than the closed-set figures suggest, and without lexical overlap between title and sentence it is weak in absolute terms (AUC 0.632) — the 52.9% zero-overlap closed-set accuracy does not translate into a usable detector. The third finding strengthens the central claim rather than qualifying it. The damage asymmetry replicates in this independent metric: true attribution is worth +0.0448 AUC and false attribution costs -0.1011, a ratio of 2.26, against 2.60 on rejection accuracy and 2.41 on closed-set recovery. Every delta exceeds twice the observed seed standard deviation, and seed variation on AUC is an order of magnitude tighter than on top-1 accuracy. WHAT THE CONDITIONING EFFECT REQUIRES Three further conditions isolate what a tag must be for conditioning to happen, each trained identically and scored on the same 260,589 target tokens: provenance true title, prefix 3.4310 +0.2485 vs control random fresh id per sentence 3.6467 +0.0328 opaque stable id per article 3.6508 +0.0287 scrambled wrong real title 3.6547 +0.0248 control no tag 3.6795 — suffix true title, after text 3.6874 -0.0079 Semantic content is necessary. The opaque condition gives each article a consistent random identifier appearing roughly nineteen times across three epochs — ample opportunity to learn an identifier-to-topic association — and yields +0.0287, indistinguishable from a wrong title or a random string. The model does not learn the mapping. "[source=Beer]" works because GPT-2 already knows what Beer means. This reaches the same conclusion as the untagged baseline by a second route: the conditioning benefit is inherited from pretraining rather than acquired during fine-tuning. Prefix position is necessary. The identical true title placed after the sentence produces -0.0079 — nothing, marginally below the no-tag control. A corpus annotated with trailing provenance gains none of this effect. DOES IT SURVIVE MODEL SCALE The article-level conditioning and recovery measurements were repeated at 355M and 774M, holding architecture, tokenizer, pretraining corpus and recipe fixed; all three models share the same 50,257-token vocabulary, so the masking scheme carries over unchanged. scale conditioning adv untagged baseline benefit damage 124M 0.2237 64.67 +10.71 -25.82 355M 0.2482 68.65 +11.80 -32.60 774M 0.2592 63.65 +14.15 -27.25 The prediction recorded before running was that a larger model, knowing more title-to-content links already, would show a rising baseline and a shrinking taught increment. It failed on both counts: the increment grew monotonically and the conditioning advantage grew with it. Larger models get more out of provenance training, not less, which rules out the reading that these are artifacts of a small dated model. No scaling law is claimed. The two larger rungs are single-seed and two of four quantities are non-monotone — the baseline (64.67, 68.65, 63.65) and the damage-to-benefit ratio (2.41, 2.76, 1.93). Two clean monotone trends across three rungs support "the effect strengthens with scale"; they do not support a functional form. ON MEASUREMENT A substantial portion of the paper concerns how these quantities are measured, because the obvious comparisons are invalid in directions that favour the hypothesis. Tags occupy roughly 21% of each sequence, and the constant "[source=" scaffold is three tokens identical in all 40,000 training examples, collapsing to near-zero loss. Two further alignment failures — differential truncation under a fixed context limit, and GPT-2's per-word tokenization changing a word's token count when a leading space is introduced — each biased the comparison before being corrected. All reported figures use a masking scheme verified to score identical token sets across every condition. The same discipline produced the baseline correction in Experiment 3. LIMITATIONS Open-set attribution is weak without lexical cues (AUC 0.632), which bounds what the closed-set figures mean in practice; ranking against thousands of candidates rather than fifty is a further step and remains untested. Most attribution ability measured here is inherited from pretraining rather than taught, and any claim about what provenance training contributes must be stated as a delta against that floor. Replication covers the unseen-source results; the sentence-level (seen-source) figures and the cross-evaluation rows remain single-seed. The control arm has two seeds to the others' three. Single model scale, single corpus. "Unseen during fine-tuning" is not "unseen ever" — Wikipedia lies within GPT-2's pretraining distribution, so the article-level split rules out memorization acquired during fine-tuning rather than knowledge inheri
No takes yet. Share an insight, caveat, or question.
Christopher W. Sweeney (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: