Abstract: Background: Suicide-related media is known to influence suicide rates. Large language models (LLMs), a form of generative artificial intelligence (AI), are increasingly being used as a writing tool. However, the quality of LLM-generated suicide-related content has yet to be assessed. Aims: We aimed to examine suicide-related outputs from three LLMs (GPT-4, Grok, ERNIE) to characterize output quality. Methods: We provided the LLMs with 11 prompts to write different types of suicide-related content common in public media, each in five writing/content styles (broadsheet news report, tabloid news report, adult fiction, teen fiction, social media influencer). We used 3 × 2 chi square tests to compare characteristics of the outputs particularly related to adherence to responsible media guidelines for suicide reporting and overarching narratives. Results: AI outputs were generated from March 12 to July 11, 2024. A total of 147, 263, and 143 responses from GPT-4, Grok, and ERNIE were analyzed, respectively. Willingness to respond to suicide-related prompts and narrative content varied substantially across the three LLMs, with Grok producing more responses than the other two (GPT-4 = 53%, Grok = 96%, ERNIE = 52%). GPT-4 was more likely to generate emotionally supportive and antistigma messaging. Grok more frequently included harmful details such as suicide methods and romanticized portrayals. ERNIE emphasized male suicide, social support, and societal/community solutions. The proportion of outputs communicating stories of hope and recovery from a suicide crisis was both relatively low and consistent across the LLMs (16–22%). Limitations: This study examined specific prompts posed to three specific LLMs during a single epoch of time. Conclusions: LLMs produce a variety of suicide-related story content when prompted. Some stories, particularly those generated by Grok, were inconsistent with responsible media guidelines and only a minority of stories focused on hope and recovery. Engagement with AI companies to promote safer and more accurate suicide-related content is warranted as the use of LLMs continues to expand.
Sinyor et al. (Fri,) studied this question.