PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 13, 20250 citationsOpen Access

Combating Hallucinations in Large Language Models: A Multi-Scale Dialogue Semantic Modeling Framework for Automated Depression Risk Assessment

View Full Paper
BYBo YuZRZehua RenYLYue Li

Key Points

  • The proposed framework improves depression risk assessment by reducing hallucinations in large language models.
  • Evaluation on the DAIC-WOZ dataset showed a significant correlation (Pearson r = 0.749) with actual PHQ-8 scores.
  • Mean absolute error for the development set was 3.1, indicating strong predictive accuracy for depression risk.
  • Multi-scale semantic modeling effectively captures local and broader context in dialogue data, enhancing response reliability.

Abstract

Current large language model (LLM) approaches for depression detection, which generate response vectors from prompts, often yield transcribed text that is informationally incomplete and semantically ambiguous. This frequently results in responses that seem superficially plausible yet are factually incorrect due to hallucinated reasoning. As a consequence, response vectors become contaminated with spurious information, compromising the reliability of detection outcomes. To address the challenge of hallucination in LLMs particularly in contexts with scarce conversational history or pronounced semantic ambiguity this paper introduces a novel multi-scale semantic modeling algorithm based on question-answering dialogues. The proposed method aims to support fully automated processing of dialogue data for individual depression risk prediction. Our algorithm constructs semantic representations at multiple scales using Q&A dialogue data. The first scale captures local semantics within a single dialogue turn, while subsequent scales incorporate context across two consecutive turns to model broader discourse information. Integrated with a tailored neural network architecture, the framework extracts semantic features indicative of depression risk. The methodology was evaluated experimentally using the DAIC-WOZ dataset. Results indicated a strong correlation between the screening outcomes of our depression risk assessment algorithm and actual PHQ-8 scores on the development set (Pearson r = 0.749, p < 0.05). In terms of predictive accuracy, the development set achieved a mean absolute error (MAE) of 3.1 and a root mean square error (RMSE) of 4.1, while the test set obtained an MAE of 4.12 and an RMSE of 4.79.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yu et al. (2025) studied this question.

synapsesocial.com/papers/68ed1896f29694dd1da78dffhttps://doi.org/10.22541/au.176027777.74651227/v1
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Large language models for depression recognition in spoken language integrating psychological knowledge2025 · 13 citations
  2. 2MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models2026 · 1 citations
  3. 3Depression Detection and Analysis using Large Language Models on Textual and Audio-Visual Modalities2024 · 5 citations
  4. 4The use of large language models in automated depression detection2026
  5. 5The Role of Humanization and Robustness of Large Language Models in Conversational Artificial Intelligence for Individuals With Depression: A Critical Analysis2024 · 52 citations