Benchmark success in grounded planning can be an illusion. We demonstrate that agents can achieve perfect task performance while remaining representationally blind to the objects they manipulate: a failure mode we term the Archive Dichotomy that remains invisible to standard evaluation. By forcing agents to perform explicit internal object commitment, we isolate the accumulated cost of this representational blindness, which we define as grounding debt. Our findings reveal that planning and grounding are orthogonal competencies: the same model can achieve 100% multi-step planning success while failing completely (0%) at object selection as categorical entropy scales beyond 30 objects. We provide a quantitative diagnostic framework with four measurable axes and precise empirical thresholds, including a performance cliff at 300 homogeneous trajectories and an architectural ceiling robust to 220× capacity scaling. These thresholds offer engineers a concrete methodology for measuring representational robustness, turning the illusion of benchmark performance into evidence-based architectural diagnosis.
Shoryavardhaan Gupta (Sun,) studied this question.