PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 2025Frontiers in Artificial Intelligence106 citationsOpen Access

Survey and analysis of hallucinations in large language models: attribution to prompting strategies or model behavior

View Full Paper
DADang Anh-HoangVTVu TranLNLe-Minh Nguyen

Key Points

  • Hallucinations in large language models can be influenced by both prompting strategies and intrinsic model behavior, leading to varying outputs.
  • The study evaluates several state-of-the-art models like GPT-4 and LLaMA 2 using benchmarks for factuality assessment.
  • The authors introduce a framework that quantifies the role of Prompt Sensitivity and Model Variability in hallucination attributions.
  • Findings highlight effective prompting strategies like chain-of-thought prompting in reducing hallucinations while also noting model limitations.

Abstract

Hallucination in Large Language Models (LLMs) refers to outputs that appear fluent and coherent but are factually incorrect, logically inconsistent, or entirely fabricated. As LLMs are increasingly deployed in education, healthcare, law, and scientific research, understanding and mitigating hallucinations has become critical. In this work, we present a comprehensive survey and empirical analysis of hallucination attribution in LLMs. Introducing a novel framework to determine whether a given hallucination stems from not optimize prompting or the model's intrinsic behavior. We evaluate state-of-the-art LLMs—including GPT-4, LLaMA 2, DeepSeek, and others—under various controlled prompting conditions, using established benchmarks (TruthfulQA, HallucinationEval) to judge factuality. Our attribution framework defines metrics for Prompt Sensitivity (PS) and Model Variability (MV) , which together quantify the contribution of prompts vs. model-internal factors to hallucinations. Through extensive experiments and comparative analyses, we identify distinct patterns in hallucination occurrence, severity, and mitigation across models. Notably, structured prompt strategies such as chain-of-thought (CoT) prompting significantly reduce hallucinations in prompt-sensitive scenarios, though intrinsic model limitations persist in some cases. These findings contribute to a deeper understanding of LLM reliability and provide insights for prompt engineers, model developers, and AI practitioners. We further propose best practices and future directions to reduce hallucinations in both prompt design and model development pipelines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Anh-Hoang et al. (2025) studied this question.

synapsesocial.com/papers/68de84b65b556a9128e1b460https://doi.org/10.3389/frai.2025.1622292
Ask AI
Helpful
Bookmark
Share
View Full Paper