PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20251 citationsOpen Access

Can LLMs Explain Themselves Counterfactually?

View Full Paper
ZDZahra DehghanighobadiAFAsja FischerMZMuhammad Bilal Zafar

Key Points

  • Self-generated counterfactual explanations from large language models often lack accuracy and coherence.
  • Testing across various model families highlighted inconsistencies in generating counterfactual outputs.
  • Model temperature settings impact the quality of explanations, indicating a nuanced approach is necessary.
  • These findings suggest the need for improved methods to enhance self-explanatory features in language models.

Abstract

Explanations are an important tool for gaining insights into the behavior of ML models, calibrating user trust and ensuring regulatory compliance. Past few years have seen a flurry of post-hoc methods for generating model explanations, many of which involve computing model gradients or solving specially designed optimization problems. However, owing to the remarkable reasoning abilities of Large Language Model (LLMs), self-explanation, that is, prompting the model to explain its outputs has recently emerged as a new paradigm. In this work, we study a specific type of self-explanations, self-generated counterfactual explanations (SCEs). We design tests for measuring the efficacy of LLMs in generating SCEs. Analysis over various LLM families, model sizes, temperature settings, and datasets reveals that LLMs sometimes struggle to generate SCEs. Even when they do, their prediction often does not agree with their own counterfactual reasoning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dehghanighobadi et al. (2025) studied this question.

synapsesocial.com/papers/68f0d5eb105731330a2b2155https://doi.org/10.48550/arxiv.2502.18156
Ask AI
Helpful
Bookmark
Share
View Full Paper