Key points are not available for this paper at this time.
Enterprise strategic decisions must reconcile conflicting stakeholder interests under time pressure, yet the consulting that traditionally supports them remains out of reach for most small and medium-sized enterprises. Large language models make parts of this work automatable, but a single model speaks with one voice, and its reasoning can be neither inspected nor contested. As frontier models continue to improve, whether structured multi-agent deliberation is worth its additional cost has therefore become an empirical question rather than a design assumption. This study designs, prototypes, and evaluates Servi.AI, a personality-aware multi-agent intelligent decision support system for enterprise strategy. The system grounds every recommendation in a traceable evidence chain retrieved over a knowledge graph. It stages a statement–discussion–consensus roundtable in which role-specialized agents argue from conflicting professional stances. It also simulates how synthetic stakeholders, calibrated against a public personality dataset of 874,434 respondents, will experience the candidate decision. The roundtable characterizes how a decision is argued, whereas the sandbox characterizes how it will be experienced. A questionnaire with 133 screened decision-makers confirms these requirement priorities. We evaluate the system across five experimental axes: 2800 controlled simulation runs and a twelve-case benchmark judged blind across three model families. A strong single model attains the highest holistic scores (8.22–8.56/10 across judges), while deliberation contributes auditable role-grounded reasoning, conflict surfacing, and an executable blueprint. Ablating retrieval loses all 71 valid pairwise comparisons, while a heterogeneous five-family agent pool significantly improves risk coverage (Cliff’s δ=+0.75). Retrieval thus drives the evidence-side qualities, and the role structure drives the deliberation-side ones: deliberative value is decomposable along architectural components, a middle-range design proposition. These findings support selective rather than default deployment. The released benchmark, judging protocol, and raw results provide a reusable basis for deciding when multi-agent decision support is worth its cost.
Zhou et al. (2026) studied this question.