Computational evaluation uncovers strict preconditions for typed decay to alter retrieval outcomes, demonstrating that standard benchmarks rarely satisfy requirements for performance gains.
We give a computable, retriever-independent precondition for whether retrieval-time decay can change an outcome, and reanalyse a contradictory literature. Under a score that composes relevance with salience, per-class half-lives move a top-k result only where the candidate set mixes retention classes: inside one class the salience order is the recency order for every finite half-life, so typing does nothing. It is checkable before the system it would evaluate exists: the certified fraction never exceeds a straddle rate from timestamps alone. The field's benchmarks fail it: of fifteen corpus/task entries from three benchmarks, twelve certify at exactly 0.000 under a multi-class age typing, and the median candidate span separates them at the rule's threshold. Where the mechanism can act, it is strong: on a corpus built to contain the failure it corrects, typed decay takes accuracy from 0.500 to 0.975 and changes nothing in a single-class span under that scoring. Both halves are demonstrated causally on an independent retriever that is deliberately not our store, and the win is compositional — not a claim that decay alone moves outcomes. Outside it we find no win on any whole corpus the instrument admitted one on — including the one it nominates as most able, where typed decay loses at every threshold. One slice runs the other way, on the exact failure the corrective was built for, in a build with no relevance signal. On the one corpus meeting all three preconditions with classes a world model declares, all three arms return precision@1 0.250 — a tie isolating a fourth condition: the salience ratio the workload's ages supply must exceed the relevance ratio the answer must beat. The contradictions are this precondition, unstated: three conditions are readable from a corpus before any system exists, the fourth needs golds and a relevance model.
No takes yet. Share an insight, caveat, or question.
Friedemann Lipphardt (2026) studied this question.