Software evaluation demonstrates deterministically verifiable agent claims using artifact provenance, suggesting robust alternatives to non-deterministic model graders.
An LLM grader may still disagree with itself when temperature=0. Earlier work therefore measured seed settings, repeated epochs, and grade variance instead of treating a single verdict as deterministic. For long-running agents, some claims allow a simpler approach: the final verdict can be computed from process and artifact evidence without giving the model that authority. Independent Evidence Resolution Layer 1 (IERL-1) is a clean-room reference implementation for this class of claim. It checks a hash-chained journal, a process identity bound by the controller, exit status, artifact SHA-256, policy, semantic conditions, and event order. Model text remains an observation unless the required evidence supports promotion to VERIFIED. The verifier is fallible too. Five automated adversarial reviews found twelve defects after earlier checks had passed; each repair was followed by a new evidence run, and the superseded run was kept as a record of the earlier apparatus. The v1.5 snapshot passes 25 unit tests, 57 conformance executions, eight apparatus probes, and a separately implemented auditor within the stated test matrix. IERL-1 is a reference method, not a Codex patch or formal proof. Hash chains, provenance, independent verification, and fail-closed validation are established techniques. IERL-1 uses them to control when an execution claim may enter later context as verified, with the verifier kept open to later counterexamples.
No takes yet. Share an insight, caveat, or question.
H. Tamba (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: