Abstract AI systems increasingly produce outputs presented as physical discoveries: recovered equations, learned causal structures, transferable representations. Yet the inference from instrumental success to physical applicability is rarely made explicit, and its logical structure is almost never examined. This paper introduces a claim-level diagnostic verdict protocol that separates formal or computational success from physical applicability for AI-based claims about the physical world. The protocol takes as input a claim text, its intended regime, data pipeline, and evaluation protocol, and returns a structured output: a verdict (PASS, Level-of-Applicability Concern LoA, or OPEN), a trigger report identifying which applicability conditions are at issue, and a reformulation directive specifying how the claim can be rewritten into an assessable, regime-bound form. Three domain-specific triggers, derived as specializations of the general applicability conditions A1–A6 developed in companion work, operationalize the protocol: T-VG (Vocabulary Guard) halts interpretive upgrades from performance to discovery language; T-SC (Synthetic Confirmation) flags cases where training success cannot serve as its own warrant because the architecture encodes the target properties; T-OOD (Regime-Stability) requires demonstrated stability under physically meaningful regime variations. Three additional stop-rules (SR1–SR3) enforce governance against interpretive drift across the protocol. We apply the protocol to three cases with distinct verdict profiles: a foundation-model weather forecast (GraphCast; Lam et al. 2023), a symbolic-regression equation-discovery system (AI Feynman; Udrescu and Tegmark 2020), and a multi-physics transfer-learning claim (McCabe et al. 2023). The protocol does not decide whether AI can do physics; it provides the reporting discipline required for AI-based physics claims to become assessable. Keywords: AI in physics; foundation models; equation discovery; applicability conditions; levels of description; epistemic opacity; synthetic confirmation; regime specification; interpretive drift; diagnostic protocol; scientific inference; scientific machine learning
Zierhut et al. (2026) studied this question.