This work presents a governance-driven approach to sim-to-sim residual learning in robotics, addressing the failure modes caused by misaligned paired transitions across simulators. We introduce a deterministic pairing protocol that enforces strict episode and timestep correspondence between simulators, combined with a hard alignment gate that blocks training when pairing validity fails. A projection-consistent residual correction model is then applied to correct next-state discrepancies while preserving valid state geometry. Evaluation is performed using a horizon-validated protocol, including one-step accuracy, teacher-forced rollouts at multiple horizons (50/200/500), contact-regime slices, and free-running stability checks. The results demonstrate that residual correction remains stable and accurate under contact-rich conditions when governed by explicit pairing and validation constraints. This repository contains the reference implementation, datasets, and validation artifacts for reproducing the reported results.
Olevester Anthony (Sun,) studied this question.