Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper) | Synapse