PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 27, 2026Strategic Management Journal0 citationsOpen Access

Bias in, symbolic compliance out? GPT 's reliance on gender and race in strategic evaluations

View Full Paper
TBTristan L. BotelhoQWQingyang (Iris) WangYale University

Key Points

  • The aim is to investigate how large language models, like GPT, rely on gender and race in evaluating startup pitches.
  • Conducted 26,000 evaluations of identical startups with varying founder names to simulate gender and racial perceptions.
  • Performed 'Second Opinion' experiments where GPT evaluated pitches alongside simulated human bias to analyze bias correction mechanisms.
  • GPT did not systematically assign lower scores to underrepresented minorities but avoided ranking them last.
  • Corrections to explicit identity-based bias were more readily made than to biases presented as neutral critiques, though limited in extent.

Abstract

Abstract Research summary Organizations are increasingly using large language models (LLMs) to support strategic evaluations. We examine whether and how these systems rely on gender and race. We asked GPT to evaluate identical startup pitches varying only the founder's name, shaping gender and race perceptions. Across 26,000 evaluations, GPT did not systematically assign lower scores to underrepresented minorities but avoided ranking them last without increasing winning likelihoods. To explain these patterns, we conducted “Second Opinion” experiments where GPT evaluated pitches alongside inputs simulating human bias. GPT more readily corrected explicit, identity‐based bias than bias framed as neutral business critiques, with corrections limited in magnitude. We theorize these findings reflect symbolic compliance : LLMs suppress overt discrimination without substantively altering evaluative logic, allowing inequality to persist in AI‐supported strategic evaluations. Managerial summary Large language models (LLMs), like OpenAI's ChatGPT, are increasingly used in strategic evaluations (e.g., hiring, pitches). We examine whether and how these models exhibit gender and racial biases in their evaluations of startup pitches, where we only varied founder names (shaping gender and race perceptions). Across multiple experiments, we find that GPT evaluators did not systematically assign lower scores to underrepresented minorities, primarily by reducing their likelihood of being ranked last. However, this behavior reflects a symbolic effort to avoid overt discrimination rather than a deeper fairness commitment. While LLMs may not reproduce historical and societal biases in overt form, their ability to correct them remains limited. These results highlight the need for implementing bias mitigation measures before integrating LLMs into high‐stakes strategic evaluation processes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Botelho et al. (2026) studied this question.

synapsesocial.com/papers/69eefe1efede9185760d4c18https://doi.org/10.1002/smj.70094
Ask AI
Helpful
Bookmark
Share
View Full Paper