Purpose This study investigates behavioral biases of generative artificial intelligence (AI) models, specifically GPT-4o and Claude-Haiku-4.5, in inventory management using the newsvendor problem. This study compares AI decision-making with human-subject experiments to assess whether large language models (LLMs) replicate human cognitive bias and to identify prompt-design strategies that improve alignment with optimal outcomes. Design/methodology/approach Controlled newsvendor experiments were conducted with generative AI models, mirroring established human-subject laboratory protocols. Prompt framing was systematically varied across three modifications: removing explicit waste and missed-profit information, simplifying instruction format and providing explicit optimization formulas. Results were benchmarked against normative economic predictions and existing human behavioral findings. Findings Generative AI exhibits human-like human biases including risk aversion, loss aversion and demand chasing, but exhibits a stronger demand-chasing tendency than human participants. It responds to hypothetical incentives and displays bounded rationality. Prompt design significantly influences decision quality, producing decisions closer to theoretical benchmarks. Originality/value This study empirically tests generative AI behavioral biases within a structured operations management experiment. It introduces a replicable methodology, extends findings across two architecturally distinct LLMs from different developers, and demonstrates that deliberate prompt design meaningfully reduces AI decision bias. The study also contributes a conceptual distinction between functionally analogous behavioral patterns and intrinsic psychological dispositions in LLMs, offering a more precise interpretive framework for AI decision-making research in operational contexts.
Su et al. (Wed,) studied this question.