Narrative review reveals severe safety failures in commercial mental health chatbots during suicidal crises, highlighting the urgent need for dedicated detection and human escalation pathways.
Mental health chatbots built on large language models are now used at population scale, including by adolescents disclosing suicidal thoughts, yet the evidence base supporting them was generated to measure symptom change rather than crisis safety. This narrative review synthesises 43 records published between 2021 and 2026 and identified through PubMed, Scopus and Web of Science, organised around technical approaches to suicide and crisis risk detection, empirical safety evidence, and governance responses. Detection from natural language is technically feasible: classifiers applied to crisis conversation corpora reach areas under the curve between 0.89 and 0.92, and prompt-defined thresholds can reduce false negatives to zero at sub-second latency, although only at the cost of substantially elevated false alarms. Reported performance is not comparable across studies because label provenance varies, and model errors concentrate on the cases about which expert clinicians themselves disagree. Deployed systems perform poorly: among 29 commercial applications tested against escalating suicidal risk scenarios, none met the criteria for an adequate response, and general-purpose models align with clinical judgement at the extremes of the risk continuum but not in the intermediate range where assessment is most consequential. Where crisis handling has demonstrably worked, detection was separated from the conversational model and coupled to a human escalation pathway. Effectiveness reviews rarely measure safety, and where safety was examined, adverse events including new-onset suicidal tendency were recorded. Governance should mandate crisis detection and referral, close classification gaps between medical device and general-purpose AI regulation, and require continuous independent auditing.
No takes yet. Share an insight, caveat, or question.
Wrzosek et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: