PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2024Political Analysis146 citationsOpen Access

Synthetic Replacements for Human Survey Data? The Perils of Large Language Models

View Full Paper
JBJames BisbeeJCJoshua D. ClintonCDCassy Dorff

Key Points

  • Synthetic data generated by ChatGPT shows averages resembling real surveys but lacks statistical reliability.
  • Variability in synthetic responses is less than in real surveys, as found in the American National Election Study.
  • Analysis using large language models highlights significant differences in regression coefficients compared to actual data estimates over time and prompts used. Synthetic responses can vary with slight changes in wording, suggesting instability in output.

Abstract

Abstract Large language models (LLMs) offer new research possibilities for social scientists, but their potential as “synthetic data” is still largely unknown. In this paper, we investigate how accurately the popular LLM ChatGPT can recover public opinion, prompting the LLM to adopt different “personas” and then provide feeling thermometer scores for 11 sociopolitical groups. The average scores generated by ChatGPT correspond closely to the averages in our baseline survey, the 2016–2020 American National Election Study (ANES). Nevertheless, sampling by ChatGPT is not reliable for statistical inference: there is less variation in responses than in the real surveys, and regression coefficients often differ significantly from equivalent estimates obtained using ANES data. We also document how the distribution of synthetic responses varies with minor changes in prompt wording, and we show how the same prompt yields significantly different results over a 3-month period. Altogether, our findings raise serious concerns about the quality, reliability, and reproducibility of synthetic survey data generated by LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bisbee et al. (2024) studied this question.

synapsesocial.com/papers/68e69843b6db64358761e5a8https://doi.org/10.1017/pan.2024.5
Ask AI
Helpful
Bookmark
Share
View Full Paper