Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khati, Dipin, Rodriguez-Cardenas, Daniel, Palacio, David N., Velasco, Alejandro, Tufano, Michele, Poshyvanyk, Denys
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915932554657792
author Khati, Dipin
Rodriguez-Cardenas, Daniel
Palacio, David N.
Velasco, Alejandro
Tufano, Michele
Poshyvanyk, Denys
author_facet Khati, Dipin
Rodriguez-Cardenas, Daniel
Palacio, David N.
Velasco, Alejandro
Tufano, Michele
Poshyvanyk, Denys
contents As Large Language Models for Code (LM4Code) become integral to software engineering, establishing trust in their output becomes critical. However, standard accuracy metrics obscure the underlying reasoning of generative models, offering little insight into how decisions are made. Although post-hoc interpretability methods attempt to fill this gap, they often restrict explanations to local, token-level insights, which fail to provide a developer-understandable global analysis. Our work highlights the urgent need for \textbf{global, code-based} explanations that reveal how models reason across code. To support this vision, we introduce \textit{code rationales} (CodeQ), a framework that enables global interpretability by mapping token-level rationales to high-level programming categories. Aggregating thousands of these token-level explanations allows us to perform statistical analyses that expose systemic reasoning behaviors. We validate this aggregation by showing it distills a clear signal from noisy token data, reducing explanation uncertainty (Shannon entropy) by over 50%. Additionally, we find that a code generation model (\textit{codeparrot-small}) consistently favors shallow syntactic cues (e.g., \textbf{indentation}) over deeper semantic logic. Furthermore, in a user study with 37 participants, we find its reasoning is significantly misaligned with that of human developers. These findings, hidden from traditional metrics, demonstrate the importance of global interpretability techniques to foster trust in LM4Code.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16771
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
Khati, Dipin
Rodriguez-Cardenas, Daniel
Palacio, David N.
Velasco, Alejandro
Tufano, Michele
Poshyvanyk, Denys
Software Engineering
Machine Learning
As Large Language Models for Code (LM4Code) become integral to software engineering, establishing trust in their output becomes critical. However, standard accuracy metrics obscure the underlying reasoning of generative models, offering little insight into how decisions are made. Although post-hoc interpretability methods attempt to fill this gap, they often restrict explanations to local, token-level insights, which fail to provide a developer-understandable global analysis. Our work highlights the urgent need for \textbf{global, code-based} explanations that reveal how models reason across code. To support this vision, we introduce \textit{code rationales} (CodeQ), a framework that enables global interpretability by mapping token-level rationales to high-level programming categories. Aggregating thousands of these token-level explanations allows us to perform statistical analyses that expose systemic reasoning behaviors. We validate this aggregation by showing it distills a clear signal from noisy token data, reducing explanation uncertainty (Shannon entropy) by over 50%. Additionally, we find that a code generation model (\textit{codeparrot-small}) consistently favors shallow syntactic cues (e.g., \textbf{indentation}) over deeper semantic logic. Furthermore, in a user study with 37 participants, we find its reasoning is significantly misaligned with that of human developers. These findings, hidden from traditional metrics, demonstrate the importance of global interpretability techniques to foster trust in LM4Code.
title Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2503.16771