The rapid advancement of Generative AI, particularly large language models (LLMs), has sparked an extensive debate regarding the use of synthetic data in social research. Beyond the profound epistemological implications of this possibility, the first prerequisite for its feasibility is to test if such models adjust their responses to capture the vast individual and social diversity found in human populations and, if so, to what extent their outputs are comparable to those observed in human samples. This study investigates whether LLM-generated synthetic personas can accurately express basic psychological needs and human values considering a wide range of distinct individual and social positions. Additionally, to assess whether the model can maintain such consistency beyond information available within the scientific community, we propose a novel task in which the model cannot rely on prior knowledge. The findings suggest that a language model, such as GPT-4o , embodies a set of implicit theories about how a human with specific characteristics would respond to a questionnaire on basic needs and values, and it can apply these theories with sufficient consistency both when responding to an established scale and to a newly developed one. We further examine how certain aspects of our results align with key limitations identified in the critical literature on the use of synthetic data.
Chulvi et al. (Wed,) studied this question.