PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 15, 20260 citationsOpen Access

Mind the Ladder: A Benchmark for Level 1--3 Causal Reasoning in World Models

View Full Paper
DPDi Prodi Paolo

Key Points

  • To establish a benchmark for assessing levels of causal reasoning in latent world models using Pearl's Ladder of Causality.
  • Introduced Mind the Ladder as a diagnostic benchmark for latent world models.
  • Operated on three levels of causality (association, intervention, counterfactuals) in the latent space.
  • Validated on the Glitched Hue Two Room environment to test for causal disentanglement.
  • VoE surprise measure does not reliably indicate causal accuracy; many models show high surprise without passing counterfactual tests.
  • Models can have high surprise due to physical violations yet fail Level 3 counterfactual assessments.

Abstract

World models based on Joint-Embedding Predictive Architecture (JEPA) have demonstrated emergent physical understanding through Violation-of-Expectation (VoE) paradigms. However, the "surprise" metric used to evaluate these models conflates statistical novelty with genuine causal reasoning. This paper introduces Mind the Ladder, a diagnostic benchmark and metric suite for testing causal fidelity in latent world models. The framework operationalises Pearl's Ladder of Causality (Level 1: Association, Level 2: Intervention, Level 3: Counterfactuals) directly in the latent space of a trained world model, making it architecture-agnostic. Three novel metrics are proposed: AAP Surprise Ratio, Structural Invariance, and AAP Consistency Advantage all grounded in the LeWorldModel (LeWM) architecture. The benchmark is validated on the Glitched Hue Two Room environment, which tests causal disentanglement between spurious correlations and true causal mechanisms. Results show that VoE surprise alone is insufficient: a model can exhibit high surprise for physical violations while still failing Level 3 counterfactual tests.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Di Prodi Paolo (2026) studied this question.

synapsesocial.com/papers/6a06b940e7dec685947abdd7https://doi.org/10.5281/zenodo.20162155
Ask AI
Helpful
Bookmark
Share
View Full Paper