Authors
Loading...
Benchmark-driven evaluation demonstrates language model performance in simulated embodied reasoning tasks, suggesting strengths and limitations.
Demirhan et al. (2026) studied this question.