Large language models (LLMs) are increasingly used to simulate infrastructure tool behavior—predicting command outputs and state transitions for Docker, Kubernetes, Terraform, and other DevOps tools without executing them. While prior work has established that LLMs achieve approximately 60% single-step accuracy on text-based state prediction, no benchmark exists for infrastructure-specific domains, and the impact of error compounding across multi-step operational sequences remains unquantified. We present the first multi-domain infrastructure state prediction benchmark comprising 1,496 evaluation entries across 80 scenarios spanning seven infrastructure domains plus cross-domain interactions, generated from containerized real-tool execution.
Swaraj Dhondge (Mon,) studied this question.