Playing the network backward: A Game Theoretic Attribution Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zimmermann, Jakob Paul, Berend, Jim, Loho, Georg, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917468217278464
author Zimmermann, Jakob Paul
Berend, Jim
Loho, Georg
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Zimmermann, Jakob Paul
Berend, Jim
Loho, Georg
Lapuschkin, Sebastian
Samek, Wojciech
contents Attribution methods explain which input features drive a model's prediction, making them central to model debugging and mechanistic interpretability. Yet backward attribution methods, including gradients, LRP, and transformer-specific rules, lack a shared framework in which to compare the underlying backward calculations. We introduce such a framework by recasting backward attribution as a two-player game on an extended network graph, building on Gaubert and Vlassopoulos' ReLU Net Game. Gradients and the full alpha-beta-LRP family arise as integrals over game trajectories under specific equilibria, so attribution maps become projections of trajectory distributions rather than the primary object. Desired explanation properties, such as localisation focus, robustness to input noise, or stable attention routing, can be specified as game-theoretic concepts, including policy regularization, risk aversion, and extended action sets, and translate directly into novel adaptations of the well-known backward rules. On ViT-B/16, one such selected adaptation of alpha-beta-LRP outperforms prior transformer-specific backward methods across all considered localisation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06212
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Playing the network backward: A Game Theoretic Attribution Framework
Zimmermann, Jakob Paul
Berend, Jim
Loho, Georg
Lapuschkin, Sebastian
Samek, Wojciech
Machine Learning
Computer Vision and Pattern Recognition
Attribution methods explain which input features drive a model's prediction, making them central to model debugging and mechanistic interpretability. Yet backward attribution methods, including gradients, LRP, and transformer-specific rules, lack a shared framework in which to compare the underlying backward calculations. We introduce such a framework by recasting backward attribution as a two-player game on an extended network graph, building on Gaubert and Vlassopoulos' ReLU Net Game. Gradients and the full alpha-beta-LRP family arise as integrals over game trajectories under specific equilibria, so attribution maps become projections of trajectory distributions rather than the primary object. Desired explanation properties, such as localisation focus, robustness to input noise, or stable attention routing, can be specified as game-theoretic concepts, including policy regularization, risk aversion, and extended action sets, and translate directly into novel adaptations of the well-known backward rules. On ViT-B/16, one such selected adaptation of alpha-beta-LRP outperforms prior transformer-specific backward methods across all considered localisation metrics.
title Playing the network backward: A Game Theoretic Attribution Framework
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.06212