Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 24, 2026ACM Transactions on Software Engineering and MethodologyOpen Access

Toward Explaining Large Language Models in Software Engineering Tasks

View Full Paper
Ask AI
Bookmark
Share

Authors

AVAntonio VitalePolytechnic University of TurinKNKhai-Nguyen NguyenWilliam & MaryDPDenys PoshyvanykWilliam & Mary

Discussion

Loading...

Member takes

Overview

Evaluation study reveals that FeatureSHAP improves explanation fidelity and reduces computational cost in software engineering LLMs, suggesting a path to transparent AI development.

Key Points

  • To introduce and evaluate FeatureSHAP, a model-agnostic explainability framework designed to explain LLM decisions in software engineering tasks using Shapley values.
  • Designed FeatureSHAP to attribute outputs to high-level input features via systematic input perturbation and task-specific similarity metrics across open-source and proprietary LLMs.
  • Evaluated the framework on code generation and code summarization tasks alongside a practitioner survey with 37 participants.
  • FeatureSHAP assigned less importance to irrelevant features and achieved higher fidelity explanations compared with baseline methods.
  • The framework operated at substantially lower computational cost than token-level alternatives and improved decision-making among 37 surveyed software practitioners.

Cite This Study

Vitale et al. (2026) studied this question.

synapsesocial.com/papers/6a8c0010bca056c88e6dee70https://doi.org/10.1145/3837084
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization2024 · 14 citations
  2. 2Toward a Theory of Causation for Interpreting Neural Code Models2024 · 13 citations
  3. 3Efficient Memory Management for Large Language Model Serving with PagedAttention2023 · 1,590 citations
  4. 4Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models2025 · 11 citations
  5. 5A Critique and Improvement of an Evaluation Metric for Text Segmentation2002 · 434 citations