A short observation note documenting a recurring class of AI errors in visual-spatial reasoning (e.g., mirroring vs. rotating images, "flip the page" ambiguity). The note argues that such errors stem not from reasoning failure but from the absence of embodied cognition — AI models process images as pixel matrices without a physical frame of reference, unlike humans who mentally simulate physical actions. The note discusses why this gap is temporary and likely to narrow as models are trained on data with implicit spatial/physical grounding (video, robotics, multimodal datasets). Originally observed and noted in 2024; revised June 2026.
Serhii Kanivets (Sat,) studied this question.