PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 27, 2026Decision Analysis2 citations

Assessing the Quality of Large Language Models and Human Inputs to a Decision: A Proposed Framework and Two Benchmark Case Studies

View Full Paper
AAAli E. AbbasAHAndrea C. HupmanRMRachel Monger

Key Points

  • This research aims to evaluate the quality of decision analysis inputs from large language models compared to those from human groups.
  • Developed a framework for assessing decision analysis inputs from LLMs and human groups.
  • Conducted two benchmark case studies using ChatGPT version 4.0 to illustrate the framework.
  • Compared the efficacy of inputs from crowdsourcing and LLMs based on relevancy judged by a panel.
  • Panel judgements showed high correlation in relevance assessment among inputs.
  • Human groups outperformed LLMs in generating diverse alternatives.
  • LLMs excelled at identifying uncertainties with relevant inputs.
  • Both sources performed similarly in generating preferences.

Abstract

This paper proposes a general framework for assessing the quality of inputs to a decision analysis provided by large language models (LLMs). The paper provides two benchmark case studies that focus on alternatives, preferences, and uncertainties related to a decision and are used to illustrate the proposed framework using ChatGPT version 4.0. The analysis uses the proposed framework and the data obtained to compare the efficacy of decision inputs provided by crowdsourcing from a group and those obtained from LLMs, with the relevance of inputs determined independently by a panel. The results show that (i) panel judgements about the relevance of inputs exhibited high correlation to one another; (ii) human groups performed better on generating alternatives, with higher rates of relevant alternatives; (iii) LLMs performed better on generating uncertainties, with higher rates of relevant alternatives; and (iv) human groups and LLMs performed similarly on generating preferences. These findings repeated across both subsets of data. Direct questions to participants about which input source they preferred resulted in a slight edge for artificial intelligence inputs. Although the benchmark case studies used ChatGPT version 4.0, the general framework applies to any LLM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abbas et al. (2026) studied this question.

synapsesocial.com/papers/69c6207d15a0a509bde18f3ehttps://doi.org/10.1287/deca.2025.0467
Ask AI
Helpful
Bookmark
Share
View Full Paper