Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kariyappa, Sanjay, Lécué, Freddy, Mishra, Saumitra, Pond, Christopher, Magazzeni, Daniele, Veloso, Manuela
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916274375753728
author Kariyappa, Sanjay
Lécué, Freddy
Mishra, Saumitra
Pond, Christopher
Magazzeni, Daniele
Veloso, Manuela
author_facet Kariyappa, Sanjay
Lécué, Freddy
Mishra, Saumitra
Pond, Christopher
Magazzeni, Daniele
Veloso, Manuela
contents This paper proposes Progressive Inference - a framework to compute input attributions to explain the predictions of decoder-only sequence classification models. Our work is based on the insight that the classification head of a decoder-only Transformer model can be used to make intermediate predictions by evaluating them at different points in the input sequence. Due to the causal attention mechanism, these intermediate predictions only depend on the tokens seen before the inference point, allowing us to obtain the model's prediction on a masked input sub-sequence, with negligible computational overheads. We develop two methods to provide sub-sequence level attributions using this insight. First, we propose Single Pass-Progressive Inference (SP-PI), which computes attributions by taking the difference between consecutive intermediate predictions. Second, we exploit a connection with Kernel SHAP to develop Multi Pass-Progressive Inference (MP-PI). MP-PI uses intermediate predictions from multiple masked versions of the input to compute higher quality attributions. Our studies on a diverse set of models trained on text classification tasks show that SP-PI and MP-PI provide significantly better attributions compared to prior work.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02625
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
Kariyappa, Sanjay
Lécué, Freddy
Mishra, Saumitra
Pond, Christopher
Magazzeni, Daniele
Veloso, Manuela
Machine Learning
Artificial Intelligence
This paper proposes Progressive Inference - a framework to compute input attributions to explain the predictions of decoder-only sequence classification models. Our work is based on the insight that the classification head of a decoder-only Transformer model can be used to make intermediate predictions by evaluating them at different points in the input sequence. Due to the causal attention mechanism, these intermediate predictions only depend on the tokens seen before the inference point, allowing us to obtain the model's prediction on a masked input sub-sequence, with negligible computational overheads. We develop two methods to provide sub-sequence level attributions using this insight. First, we propose Single Pass-Progressive Inference (SP-PI), which computes attributions by taking the difference between consecutive intermediate predictions. Second, we exploit a connection with Kernel SHAP to develop Multi Pass-Progressive Inference (MP-PI). MP-PI uses intermediate predictions from multiple masked versions of the input to compute higher quality attributions. Our studies on a diverse set of models trained on text classification tasks show that SP-PI and MP-PI provide significantly better attributions compared to prior work.
title Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.02625