PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026Strategy Science4 citations

How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations

View Full Paper
RARyan T. AllenRMRory McDonald

Key Points

  • The aim is to assess AI's capabilities in strategic decision making using simulations as a benchmarking tool.
  • Evaluated 21 proprietary and 13 open-source LLMs using the Back Bay Battery (BBB) simulation.
  • Developed an interface for LLMs to interact with the simulation without prior training contamination.
  • Measured performance against MBA student cohorts for comparison.
  • Later models generally outperform earlier versions in the BBB simulation.
  • Reasoning-focused models from late 2024-early 2025 exceed average MBA student scores.
  • Mid-to-late 2025 models show a decline in performance, especially in managing strategic uncertainty.

Abstract

Benchmarks have helped fuel rapid progress in large language models (LLMs) across a variety of domains including math, science, dialogue, and coding. Yet no existing benchmark adequately captures the defining elements of strategic decision making: uncertainty, complexity, irreversible multiperiod moves, and delayed or noisy feedback. This gap limits our ability to assess and guide LLMs’ capabilities in strategy. We propose that established strategy teaching simulations provide an ideal benchmarking approach because (1) they approximate the essential features of real-world strategy, and (2) they do so in a controlled, replicable environment suitable for evaluation. To demonstrate this, we assess the performance of 21 proprietary and 13 open-source LLMs on the Back Bay Battery (BBB) simulation, a widely used exercise in strategy and innovation courses. The simulation requires balancing short-term profitability against long-term competitive positioning while integrating complex information about customer preferences and technological change. We built an interface enabling LLMs to interact with the simulation as though encountering it for the first time, masking identifiers to reduce contamination from prior training data. Our results show clear progress in composite BBB performance: Later models generally outperform earlier versions, and reasoning-focused models from late 2024–early 2025 (e.g., o4-mini, Claude Sonnet 4, Gemini 2.0 Flash) exceed even the average scores of historical MBA student cohorts. However, frontier models from mid-to-late 2025 (e.g., GPT-5, Claude Opus 4.5, Gemini 3) have declined, underperforming both earlier LLMs and MBA students. This decline is partially explained by a systematic bias toward exploiting the core business at the expense of investing in future growth. Overall, these findings highlight impressive advances in LLMs’ strategic abilities since their inception. At the same time, we document current frontier models’ surprising weakness in managing strategic uncertainty. This paper pioneers and provides guidance for using simulation-based benchmarking as a productive framework for strategy researchers to track progress, identify blind spots, and shape the trajectory of strategy-specific LLM capabilities. History: Accepted for the Special Issue: Can AI Do Strategy? Supplemental Material: The online appendix is available at https://doi.org/10.1287/stsc.2025.0444 .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Allen et al. (2026) studied this question.

synapsesocial.com/papers/69b4ad7918185d8a39800b9ehttps://doi.org/10.1287/stsc.2025.0444
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Mastering the game of Go with deep neural networks and tree search2016 · 16,017 citations
  2. 2Expertise versus Bias in Evaluation: Evidence from the NIH2017 · 172 citations
  3. 3The Structure of "Unstructured" Decision Processes1976 · 3,764 citations
  4. 4POLITICS OF STRATEGIC DECISION MAKING IN HIGH-VELOCITY ENVIRONMENTS: TOWARD A MIDRANGE THEORY.1988 · 1,759 citations
  5. 5Collaborative Brokerage, Generative Creativity, and Creative Success2007 · 1,266 citations