Algorithmic audit reveals descriptor-indexed thresholds reduce prompt-channel demographic bias in medical vision-language models, indicating prompt sanitization remains essential for free-text inputs.
Part V of a series on equity and reliability in medical imaging AI. Part III of this series proved that the demographic perturbation in a medical vision–language model's prompt is exactly rank one in the standard pair readout, and drew a corollary that organised the three papers since: because the shift follows the stated descriptor rather than the patient's true group, a group-specific threshold cannot cancel it, and prompt-channel bias therefore sits upstream of every decision-rule remedy. The proof is correct. The corollary was read too broadly, including by us. A threshold indexed on the stated descriptor is a different object, and the deployer knows that descriptor exactly, having written the prompt. Two such thresholds work on all seven encoders tested, at exactly zero AUC cost, because a threshold cannot change AUC: Matching the descriptor's calibration false-negative rate to the neutral prompt's removes 79.6% to 95.2% of the deviation from the neutral decision (TITAN: 79.6%, 95% CI 74.1–80.1, patient-clustered bootstrap over 6,582 positive patients; 66.0% of decision changes). A closed-form offset of the threshold by ⟨ē, δd⟩, which needs no labels at all, removes 15.0% to 83.7% (TITAN: 60.5%, CI 55.4–62.5). Across BiomedCLIP, PubMedCLIP, OpenAI CLIP, PLIP, QuiltNet-B-32, CONCH, and the slide-level TITAN — two imaging domains, patch and whole-slide granularity — the FNR-matched threshold is the only intervention in this series that removes most of the effect on every model tested. The statistical footing is narrower than “seven encoders” suggests, and the paper says where. Four of the seven carry patient identifiers and therefore patient-clustered intervals (TITAN and the three radiology models, 10,000 replicates each); the three pathology patch encoders do not, because the public CRC-VAL-HE-7K mirror records none. On the four testable encoders every interval excludes zero, but the FNR-matched threshold's margin over the closed-form offset is confirmed on only two: PubMedCLIP's paired interval spans −2.66 to +31.81 over three findings, and OpenAI CLIP's −2.72 to +5.14 over two findings and 198 positive patients. The three encoders showing the largest apparent margins (CONCH, +73 points) are precisely the three that admit no interval. So the ordering is a direction that holds on seven models and a magnitude established on two. And both fail in the regime that motivates them. Each needs the target finding's neutral prompt, so each is available exactly when prompt sanitisation is available — and sanitisation is better, giving zero disparity at zero cost, exactly. The case for a correction rests on deployments where the descriptor cannot be removed: free-text prompts, upstream-assembled context. There the correction must be precomputed and transferred across findings, and all three such variants fail a pre-stated 80% bar — the per-image correction table at 49.8%, the transferred offset at 38.6% (CI 34.6–41.3), the transferred FNR-match at 38.5% (CI 34.3–40.9), the last two statistically indistinguishable (difference −0.33 pp, CI −1.24 to +0.57) and both negative on two of seven encoders. A descriptor's perturbation is only 0.383-correlated across findings, which is the measured cause. Three further approaches were eliminated, each against a criterion stated before the run, because they bound what is achievable in the readout rather than at the threshold. Projecting the descriptor subspace out of the readout is exact and free on TITAN — and generalises to descriptors absent from the entire lexicon — but costs 0.037–0.130 AUC on the three patch encoders while memorising its own basis on two of them. Restricting the projection to signal-free directions removes 15% of the perturbation. Marginalising the scoring vector over the descriptor bank gives exact invariance and fails a 0.02 AUC bar on two of four encoders. Two explanations we advanced for those failures also failed, refuted by conditions written into the scripts before they ran: a geometric “headroom” criterion, by CONCH holding the second-highest headroom of seven encoders and the worst measured utility cost; and a claimed impossibility theorem — that exact linear descriptor-invariance must destroy the diagnostic signal — by PubMedCLIP, where 73.7% of the signal survives such a projector. Honest scope, including about this paper's own drafts. Part V is not preregistered; PROTOCOL_PART5.md is retrospective and says so in its first line. It records, in §6, that the manuscript was revised against an adversarial re-analysis of its own first draft, which overturned three of that draft's claims: it recommended the weaker of the two corrections, compared it against an oracle for the wrong objective (Youden's J, when the endpoint is FNR disparity), reported an offset computed with information its own argument had declared unavailable, and asserted a numerical precondition (cos > 0.96) that six cohort variants cannot calibrate. It records, in §7, that a second revision followed external criticism and showed a claimed data limitation was an omission: the radiology intervals had been described as unavailable when they had merely never been computed, and computing them weakened the paper's own comparison. The three part5_review*.py scripts in the artifact contain the first re-analysis, including the tests the draft failed. The track record cuts two ways and the paper states both. Four claims across five papers have now failed under testing while every measured effect held. That is consistent with rapid self-correction — each failure was found by a falsifier written before the test ran, and reported rather than buried. It is equally consistent with publishing before conclusions have been adequately stress-tested: none of the four was caught by a reader, because these papers are self-archived, unreviewed, single-author, and unreplicated by anyone other than the author, and a referee would plausibly have caught the wrong-oracle comparison before publication rather than after. §6.3 of the manuscript notes that the two produce the same observable trace and that the author cannot adjudicate between them from inside the series, and that foregrounding one's own retractions buys credibility with the same move that demonstrates fallibility. What is auditable does not depend on resolving that: the falsifiers are in the script docstrings, the seeds are stated, and the failed arms sit in the released tables beside the successful ones. The practical payoff is narrow. The standing recommendation is unchanged from three papers ago — strip the descriptor from the prompt — and a deployer who can do that gains nothing from the corrections here. What is new is a correction to the series' theory (a descriptor-indexed threshold is eligible, contrary to Part III's restated corollary), two instruments that realise it where a simpler and better option already exists, and a measured account of why no transferable version works. Whether the motivating regime is common in deployed systems has not been established; the manuscript calls that survey the most valuable follow-up it can identify and the study that would determine whether §3 has any users at all. Files. The manuscript PDF (20 pp, 3 figures); the retrospective protocol; and an artifact archive containing all 26 analysis scripts, every result table as CSV/JSON, and the vector figures. Model weights are not redistributed — TITAN and CONCH are CC-BY-NC-ND-4.0 and were obtained through the gated Hugging Face process, which also limits who can replicate the pathology arm. Ethics. No patient data was collected. TCGA is a public consortium dataset; NIH ChestX-ray14 and CRC-VAL-HE-7K are public and de-identified. This is an audit of a model property, not a clinical study, and makes no claim about patient outcomes.
No takes yet. Share an insight, caveat, or question.
Omar Mohammed (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: