Experiments revealed that large language models exhibit some empathy structure, but lower abilities than humans, indicating room for improvement.
Empathy, a key component of human-social interaction, has become a core con-cern in human-computer interaction. This study examines whether current large language models (LLMs) can exhibit empathy in both cognitive and affective dimensions as humans. In our study, we used the standardized questionnaire to assess LLMs empathy ability and a novel paradigm was developed for LLMs eval-uation. Four main experiments were reported on LLMs empathy abilities using the Interpersonal Reactivity Index (IRI) and the Basic Empathy Scale (BES) on GPT-4 and Llama3 respectively. Two levels of evaluations were conducted to investigate whether the structural validity of the questionnaire in LLMs was aligned with humans and to compare the LLMs' empathy abilities with humans. We found GPT-4 show identical empathy dimension structure with humans while exhibiting significantly lower empathy abilities as compared to humans. Moreover, systemati-cal difference empathy ability was evident in Llama3 showing its failure to exhibit the same empathy dimensions as humans. All these findings indicate that though GPT-4 kept the same structure of human empathy (cognitive and affective), the current LLMs can not simulate empathy as we humans as indexed by the response to the questionnaire. This highlights the urgent requirements for further improving LLMs’ empathy abilities for more user-friendly human-LLMs interactions. In addition, the pipeline to generate diverse LLMs-simulated participants was also discussed.
No takes yet. Share an insight, caveat, or question.
Yu et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: