The capabilities of Large Language Models (LLMs) have rapidly evolved, enabling them to perform increasingly complex reasoning tasks. However, while their general reasoning abilities are shaped during large-scale pretraining, domain-specific reasoning such as tactical decision-making in military contexts requires dedicated post-training. This paper introduces a simulation-guided Reinforcement Fine-Tuning (RFT) approach in which reward signals are derived from the outcomes of a combat simulation environment. By embedding a military-grade simulator into the RFT loop via the Group Relative Policy Optimization (GRPO) algorithm, model outputs are evaluated based on tactical effectiveness rather than human annotations or rule-based correctness. A proof-of-concept study demonstrates that the method significantly improves the feasibility and tactical quality of generated Courses of Action (COAs), even under limited training schedules. These findings establish simulation-guided RFT as a promising direction for equipping LLMs with tactically relevant reasoning skills and pave the way toward next-generation decision-support systems in military environments.
Becker et al. (Tue,) studied this question.