Anticipating how objects will change state during ongoing activities is essential for robust video understanding and for deploying assistive, AR, and robotics systems in the real world. We introduce Object State Change Anticipation (OSCA), the problem of predicting the imminent state change of an object at the onset of the next, yet unobserved, action in egocentric procedural videos. To facilitate research on this task, we present Ego4D-OSCA, a curated benchmark that augments Ego4D segments with consistent pre-/post-state labels, inverse state relations, and a “no state change” category. Building on this dataset, we propose the first baseline for OSCA that fuses recent egocentric visual context with a structured linguistic history of actions and object states, enabling temporally grounded reasoning about near-future transformations. We establish comprehensive evaluation protocols and report extensive experiments, including ablations on modality contributions, temporal context, and robustness to class imbalance and long-tail distributions. Results demonstrate that (i) integrating structured action/state history significantly improves anticipation over visual-only variants, (ii) modeling inverse state relations benefits generalization across diverse objects and contexts, and (iii) OSCA is complementary to action anticipation and next-active object prediction, targeting a distinct predictive capability. We release Ego4D-OSCA together with code, baseline models, and reproducible training scripts to facilitate fair comparison and future advances.
Manousaki et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: