PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization

View Full Paper
SZSiyuan ZhangYZYichi ZhangYDYinpeng Dong

Key Points

  • PKUE significantly enhances the ability of large language models to manage factual hallucinations effectively, improving their accuracy.
  • The introduction of the FactualBench dataset allows for comprehensive evaluation of factual knowledge across various domains.
  • Extensive experiments show that the approach enhances performance not only in factual tasks but also in general tasks, thus broadening its applicability.
  • The findings suggest a new way to optimize knowledge use in large language models, potentially reducing the risks associated with misinformation.

Abstract

Large Language Models (LLMs) often struggle to align their responses with objective facts, resulting in the issue of factual hallucinations, which can be difficult to detect and mislead users without relevant knowledge. Although post-training techniques have been employed to mitigate the issue, existing methods usually suffer from poor generalization and trade-offs in different capabilities. In this paper, we propose to address it by directly augmenting LLM's fundamental ability to precisely leverage its knowledge and introduce PKUE, which fine-tunes the model on self-generated responses to precise and simple factual questions through preference optimization. Furthermore, we construct FactualBench, a comprehensive and precise factual QA dataset containing 181k Chinese data spanning 21 domains, to facilitate both evaluation and training. Extensive experiments demonstrate that PKUE significantly improves LLM overall performance, with consistent enhancement across factual tasks of various forms, general tasks beyond factuality, and tasks in a different language.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68da58d1c1728099cfd10e55https://doi.org/10.48550/arxiv.2502.19127
Ask AI
Helpful
Bookmark
Share
View Full Paper