An open, reproducible, motive-neutral referee for the claims that frontier AI labs and large adopters make about how much of their code an AI now writes, and for the recursion red-lines those labs state in their own safety frameworks. Eight public claims (Anthropic, OpenAI, Microsoft, Google, Meta, Salesforce, Amazon, Snap) are scored not on whether the headline percentage is true but on whether it is checkable, using four evidentiary tests that apply to any public statistic: whether the denominator is defined, whether the figure is externally audited, whether an outsider could reproduce it, and whether an independent source has checked it. Across the eight, none defines a rigorous denominator, none is externally audited, none is externally reproducible, and one carries an independent check of its specific figure. The composite score is an unweighted convenience index (mean 0.4 out of 4), not a validated metric; the individually checkable per-dimension counts are the finding. A fifth property, whether a claim isolates the research-loop subset that safety frameworks are written about, is reported but deliberately not scored, since a public statement about general coding automation has no duty to break it out; that gap is the bridge to the second layer. A second layer records the labs' own published red-lines for automated research and development (OpenAI's Preparedness Framework self-improvement thresholds and Anthropic's Responsible Scaling Policy automated-R&D thresholds) and dates their status, together with the one datable observed instance of an AI improving its own training (Google DeepMind's AlphaEvolve). No lab reports its recursion red-line crossed, and the observed instance is real but small. A dependency-free Python script (reproduce.py) recomputes the scored table, the aggregate findings and the trigger register, and exits with an error if any row is unsourced. This is a claims referee, not a capability benchmark: it complements the measurement agenda set out by Chan et al. (arXiv 2603.03992) by grading the disclosures that already exist against it, rather than proposing new metrics. Conflict of interest: the most-cited figure is Anthropic's, this work is assisted by an Anthropic model, Anthropic is also the most vocal warner about research-loop recursion, and one adopter in the set (Snap) credits Anthropic by name. Anthropic's row is marked, competitors are weighted identically, and the frontier determination is quoted directly from Anthropic's published system card. Independent analysis and open-science documentation only, not investment advice.
N Milton (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: