CafGa: Customizing Feature Attributions to Explain Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boyle, Alan, Cheng, Furui, Zouhar, Vilém, El-Assady, Mennatallah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912604470902784
author Boyle, Alan
Cheng, Furui
Zouhar, Vilém
El-Assady, Mennatallah
author_facet Boyle, Alan
Cheng, Furui
Zouhar, Vilém
El-Assady, Mennatallah
contents Feature attribution methods, such as SHAP and LIME, explain machine learning model predictions by quantifying the influence of each input component. When applying feature attributions to explain language models, a basic question is defining the interpretable components. Traditional feature attribution methods, commonly treat individual words as atomic units. This is highly computationally inefficient for long-form text and fails to capture semantic information that spans multiple words. To address this, we present CafGa, an interactive tool for generating and evaluating feature attribution explanations at customizable granularities. CafGa supports customized segmentation with user interaction and visualizes the deletion and insertion curves for explanation assessments. Through a user study involving participants of various expertise, we confirm CafGa's usefulness, particularly among LLM practitioners. Explanations created using CafGa were also perceived as more useful compared to those generated by two fully automatic baseline methods: PartitionSHAP and MExGen, suggesting the effectiveness of the system.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20901
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CafGa: Customizing Feature Attributions to Explain Language Models
Boyle, Alan
Cheng, Furui
Zouhar, Vilém
El-Assady, Mennatallah
Human-Computer Interaction
Feature attribution methods, such as SHAP and LIME, explain machine learning model predictions by quantifying the influence of each input component. When applying feature attributions to explain language models, a basic question is defining the interpretable components. Traditional feature attribution methods, commonly treat individual words as atomic units. This is highly computationally inefficient for long-form text and fails to capture semantic information that spans multiple words. To address this, we present CafGa, an interactive tool for generating and evaluating feature attribution explanations at customizable granularities. CafGa supports customized segmentation with user interaction and visualizes the deletion and insertion curves for explanation assessments. Through a user study involving participants of various expertise, we confirm CafGa's usefulness, particularly among LLM practitioners. Explanations created using CafGa were also perceived as more useful compared to those generated by two fully automatic baseline methods: PartitionSHAP and MExGen, suggesting the effectiveness of the system.
title CafGa: Customizing Feature Attributions to Explain Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2509.20901