PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 20240 citationsOpen Access

FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction

View Full Paper
ASAlessandro ScirèKGKarim GhonimRNRoberto Navigli

Key Points

Key points are not available for this paper at this time.

Abstract

Recent advancements in text summarization, particularly with the advent of Large Language Models (LLMs), have shown remarkable performance. However, a notable challenge persists as a substantial number of automatically-generated summaries exhibit factual inconsistencies, such as hallucinations. In response to this issue, various approaches for the evaluation of consistency for summarization have emerged. Yet, these newly-introduced metrics face several limitations, including lack of interpretability, focus on short document summaries (e.g., news articles), and computational impracticality, especially for LLM-based metrics. To address these shortcomings, we propose Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction (FENICE), a more interpretable and efficient factuality-oriented metric. FENICE leverages an NLI-based alignment between information in the source document and a set of atomic facts, referred to as claims, extracted from the summary. Our metric sets a new state of the art on AGGREFACT, the de-facto benchmark for factuality evaluation. Moreover, we extend our evaluation to a more challenging setting by conducting a human annotation process of long-form summarization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Scirè et al. (2024) studied this question.

synapsesocial.com/papers/68e75ddfb6db6435876d51d3https://doi.org/10.48550/arxiv.2403.02270
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Factual Consistency Evaluation of Summarisation in the Era of Large Language Models2024
  2. 2Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks2024
  3. 3Factual consistency evaluation of summarization in the Era of large language models2024 · 38 citations
  4. 4Identifying Factual Inconsistency in Summaries: Towards Effective Utilization of Large Language Model2024
  5. 5Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization2024 · 1 citations