Playing the network backward: A Game Theoretic Attribution Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Zimmermann, Jakob Paul, Berend, Jim, Loho, Georg, Lapuschkin, Sebastian, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
Hidden Monotonicity: Explaining Deep Neural Networks via their DC Decomposition
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025)
by: Pahde, Frederik, et al.
Published: (2025)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)
by: Dawoud, Karam, et al.
Published: (2023)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022)
by: Pahde, Frederik, et al.
Published: (2022)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
by: Tinauer, Christian, et al.
Published: (2024)
by: Tinauer, Christian, et al.
Published: (2024)
Iterative Inference in a Chess-Playing Neural Network
by: Sandmann, Elias, et al.
Published: (2025)
by: Sandmann, Elias, et al.
Published: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Task Attribute Distance for Few-Shot Learning: Theoretical Analysis and Applications
by: Hu, Minyang, et al.
Published: (2024)
by: Hu, Minyang, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
A comparative study on machine learning approaches for rock mass classification using drilling data
by: Hansen, Tom F., et al.
Published: (2024)
by: Hansen, Tom F., et al.
Published: (2024)
Knowledge-Guided Failure Prediction: Detecting When Object Detectors Miss Safety-Critical Objects
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Learning to Play Video Games with Intuitive Physics Priors
by: Jaiswal, Abhishek, et al.
Published: (2024)
by: Jaiswal, Abhishek, et al.
Published: (2024)
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning
by: He, Hongyang, et al.
Published: (2025)
by: He, Hongyang, et al.
Published: (2025)
PnP-Flow: Plug-and-Play Image Restoration with Flow Matching
by: Martin, Ségolène, et al.
Published: (2024)
by: Martin, Ségolène, et al.
Published: (2024)
Filtered Posterior Mean Collections: A Unified Framework for Analytical Models of Diffusion Generalization
by: Niedoba, Matthew, et al.
Published: (2026)
by: Niedoba, Matthew, et al.
Published: (2026)
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
A Unified Revisit of Temperature in Classification-Based Knowledge Distillation
by: Frank, Logan, et al.
Published: (2026)
by: Frank, Logan, et al.
Published: (2026)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Fractional Diffusion Bridge Models
by: Nobis, Gabriel, et al.
Published: (2025)
by: Nobis, Gabriel, et al.
Published: (2025)
Transforming Game Play: A Comparative Study of DCQN and DTQN Architectures in Reinforcement Learning
by: Stigall, William A.
Published: (2024)
by: Stigall, William A.
Published: (2024)
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
by: Weber, Leander, et al.
Published: (2023)
by: Weber, Leander, et al.
Published: (2023)
Drifting Fields are not Conservative
by: Franz, Leonard T., et al.
Published: (2026)
by: Franz, Leonard T., et al.
Published: (2026)
Plug and Play Active Learning for Object Detection
by: Yang, Chenhongyi, et al.
Published: (2022)
by: Yang, Chenhongyi, et al.
Published: (2022)
Enhancing Authorship Attribution with Synthetic Paintings
by: Loures, Clarissa, et al.
Published: (2026)
by: Loures, Clarissa, et al.
Published: (2026)
Attribution Upsampling should Redistribute, Not Interpolate
by: Buono, Vincenzo, et al.
Published: (2026)
by: Buono, Vincenzo, et al.
Published: (2026)
Training Feature Attribution for Vision Models
by: Bacha, Aziz, et al.
Published: (2025)
by: Bacha, Aziz, et al.
Published: (2025)
Unveiling Concept Attribution in Diffusion Models
by: Nguyen, Quang H., et al.
Published: (2024)
by: Nguyen, Quang H., et al.
Published: (2024)
Benchmarking the Attribution Quality of Vision Models
by: Hesse, Robin, et al.
Published: (2024)
by: Hesse, Robin, et al.
Published: (2024)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
Similar Items
-
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026) -
Hidden Monotonicity: Explaining Deep Neural Networks via their DC Decomposition
by: Zimmermann, Jakob Paul, et al.
Published: (2026) -
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025) -
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025) -
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)