PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 17, 20260 citationsOpen Access

Fast and Explainable Hallucination Detection for Factual QA using Constraint-Based Semantic Matching

View Full Paper
JBJyotsna Bulchandani

Key Points

  • To develop a fast and explainable method for detecting inaccuracies (hallucinations) in factual question answering.
  • Utilized a constraint-based approach focusing on named entity extraction and semantic relevance filtering.
  • Verified accuracy using length-adaptive thresholds to assess fact presence.
  • Conducted multi-domain evaluations for task-specific effectiveness.
  • Achieved 83.7% recall with 54.6% precision at 57.1% accuracy on the HaluEval QA benchmark.
  • Demonstrated 7-15× faster performance than GPT-4 with zero API costs.
  • Highlighted a 9.3% accuracy improvement from semantic filtering, reducing over-flagging.

Abstract

Large language models frequently generate plausible but factually incorrect information (hallucinations), limiting their deployment in high-stakes domains. Existing detection methods rely on expensive LLM-based judges or supervised neural models that lack interpretability. We propose a constraint-based approach for factual question answering that provides explainable, high-recall detection by extracting named entity constraints, filtering by semantic relevance, and verifying presence using length-adaptive thresholds. On the HaluEval QA benchmark, our method achieves 83.7% recall with 54.6% precision at 57.1% accuracy. While a simple entity overlap baseline achieves higher overall accuracy (65.3%), our method provides actionable constraint-level diagnostics showing which specific facts were violated, running 7-15× faster than GPT-4 with zero API costs. Ablation studies show semantic filtering contributes 9.3% accuracy by preventing over-flagging. Our high-recall, explainable approach is suitable for applications where catching hallucinations is prioritized over minimizing false alarms, such as content flagging for human review. Multi-domain evaluation reveals task-specificity: effective for factual QA but not conversational dialogue.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jyotsna Bulchandani (2026) studied this question.

synapsesocial.com/papers/696b26d7d2a12237a934a0f5https://doi.org/10.5281/zenodo.18253560
Ask AI
Helpful
Bookmark
Share
View Full Paper