AI-assisted project management tools risk generating plans and features structurally disconnected from user intent — a failure mode we term scope hallucination (syntactically valid plans semantically disconnected from validated user intent). We present a five-condition within-subject pilot study comparing a stage-gated multi-agent PM system (ProjectOS) against four LLM configurations: unguided Gemini, unguided Perplexity, unguided ChatGPT, and PM-instructed Gemini (the same model with an explicit PM system prompt). Two measures of scope discipline are assessed across two ambiguous project briefs: SMART objective quality (0–10 per criterion) and formal scope change management (4-check binary rubric). On Brief 1, where all conditions produced scoreable criteria, unguided Gemini averaged 4.25/10, Perplexity 6.5/10, and ChatGPT 5.33/10; all three scored T=0 across all 20 unguided criteria and 0/4 on scope change discipline. Adding a PM system prompt to Gemini reversed these failures: T=0 was eliminated, scope change discipline reached 4/4, and SMART quality matched ProjectOS exactly (9.5/10 each). On the strategically complex Brief 2, unguided conditions exhibited criteria avoidance entirely; PM-instructed Gemini produced criteria but at a lower average (8.25/10 vs 9.33/10 for ProjectOS), driven by the PM prompt's advisory rather than mandatory revision enforcement. Across both briefs, PM-instructed Gemini averaged 8.9/10 versus 9.4/10 for ProjectOS. These results refine the initial hypothesis: scope hallucination appears to be a property of unguided AI rather than LLMs generally. Explicit PM instruction largely eliminates the failure mode at low cost. Stage-gated architecture retains two experimentally evidenced residual advantages — a mandatory quality floor that widens with brief complexity, and embedded enforcement requiring no user configuration — as well as session-persistent state, a design property not directly tested here but unavailable in any prompt-only configuration. These are incremental rather than categorical advantages over a well-prompted LLM. The key implication for researchers is that "unguided free-form AI" is an inadequate comparator class for evaluating PM-specific AI systems.
Chintan Shelat (Thu,) studied this question.