Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
Fuente:
arXiv
Salvato in:
| Autori principali: | Bareeva, Dilyara, Yolcu, Galip Ümit, Hedström, Anna, Schmolenski, Niklas, Wiegand, Thomas, Samek, Wojciech, Lapuschkin, Sebastian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sparse, Efficient and Explainable Data Attribution with DualXDA
di: Yolcu, Galip Ümit, et al.
Pubblicazione: (2024)
di: Yolcu, Galip Ümit, et al.
Pubblicazione: (2024)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks
di: Naujoks, Jonas R., et al.
Pubblicazione: (2025)
di: Naujoks, Jonas R., et al.
Pubblicazione: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
di: Pahde, Frederik, et al.
Pubblicazione: (2025)
di: Pahde, Frederik, et al.
Pubblicazione: (2025)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
di: Vielhaben, Johanna, et al.
Pubblicazione: (2024)
di: Vielhaben, Johanna, et al.
Pubblicazione: (2024)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
di: Dreyer, Maximilian, et al.
Pubblicazione: (2025)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2025)
Finding the right XAI method -- A Guide for the Evaluation and Ranking of Explainable AI Methods in Climate Science
di: Bommer, Philine, et al.
Pubblicazione: (2023)
di: Bommer, Philine, et al.
Pubblicazione: (2023)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2024)
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
di: Puri, Bruno, et al.
Pubblicazione: (2025)
di: Puri, Bruno, et al.
Pubblicazione: (2025)
Iterative Inference in a Chess-Playing Neural Network
di: Sandmann, Elias, et al.
Pubblicazione: (2025)
di: Sandmann, Elias, et al.
Pubblicazione: (2025)
Circuit Insights: Towards Interpretability Beyond Activations
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
di: Weber, Leander, et al.
Pubblicazione: (2023)
di: Weber, Leander, et al.
Pubblicazione: (2023)
Mechanistic understanding and validation of large AI models with SemanticLens
di: Dreyer, Maximilian, et al.
Pubblicazione: (2025)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2025)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2025)
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2025)
FADE: Why Bad Descriptions Happen to Good Features
di: Puri, Bruno, et al.
Pubblicazione: (2025)
di: Puri, Bruno, et al.
Pubblicazione: (2025)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
di: Bley, Florian, et al.
Pubblicazione: (2024)
di: Bley, Florian, et al.
Pubblicazione: (2024)
Human-Centered Evaluation of XAI Methods
di: Dawoud, Karam, et al.
Pubblicazione: (2023)
di: Dawoud, Karam, et al.
Pubblicazione: (2023)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
di: Dreyer, Maximilian, et al.
Pubblicazione: (2023)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2023)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
di: Mekala, Rohan Reddy, et al.
Pubblicazione: (2024)
di: Mekala, Rohan Reddy, et al.
Pubblicazione: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
LieSolver: A PDE-constrained solver for IBVPs using Lie symmetries
di: Klausen, René P., et al.
Pubblicazione: (2025)
di: Klausen, René P., et al.
Pubblicazione: (2025)
Concept-based explanations of Segmentation and Detection models in Natural Disaster Management
di: Heydari, Samar, et al.
Pubblicazione: (2026)
di: Heydari, Samar, et al.
Pubblicazione: (2026)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
di: Becking, Daniel, et al.
Pubblicazione: (2021)
di: Becking, Daniel, et al.
Pubblicazione: (2021)
Structural Compactness as a Complementary Criterion for Explanation Quality
di: Mesgari, Mohammad Mahdi, et al.
Pubblicazione: (2026)
di: Mesgari, Mohammad Mahdi, et al.
Pubblicazione: (2026)
PINNfluence: Influence Functions for Physics-Informed Neural Networks
di: Naujoks, Jonas R., et al.
Pubblicazione: (2024)
di: Naujoks, Jonas R., et al.
Pubblicazione: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
di: Labarta, Tobias, et al.
Pubblicazione: (2026)
di: Labarta, Tobias, et al.
Pubblicazione: (2026)
Attribution-Guided Decoding
di: Komorowski, Piotr, et al.
Pubblicazione: (2025)
di: Komorowski, Piotr, et al.
Pubblicazione: (2025)
A Fresh Look at Sanity Checks for Saliency Maps
di: Hedström, Anna, et al.
Pubblicazione: (2024)
di: Hedström, Anna, et al.
Pubblicazione: (2024)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
di: Hedström, Anna, et al.
Pubblicazione: (2024)
di: Hedström, Anna, et al.
Pubblicazione: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
di: Pahde, Frederik, et al.
Pubblicazione: (2022)
di: Pahde, Frederik, et al.
Pubblicazione: (2022)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
di: Gururaj, Shreyas, et al.
Pubblicazione: (2025)
di: Gururaj, Shreyas, et al.
Pubblicazione: (2025)
Manipulating Feature Visualizations with Gradient Slingshots
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
Building Trust in PINNs: Error Estimation through Finite Difference Methods
di: Krasowski, Aleksander, et al.
Pubblicazione: (2026)
di: Krasowski, Aleksander, et al.
Pubblicazione: (2026)
CoSy: Evaluating Textual Explanations of Neurons
di: Kopf, Laura, et al.
Pubblicazione: (2024)
di: Kopf, Laura, et al.
Pubblicazione: (2024)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
di: Kahardipraja, Patrick, et al.
Pubblicazione: (2025)
di: Kahardipraja, Patrick, et al.
Pubblicazione: (2025)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Sparse, Efficient and Explainable Data Attribution with DualXDA
di: Yolcu, Galip Ümit, et al.
Pubblicazione: (2024) -
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024) -
Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks
di: Naujoks, Jonas R., et al.
Pubblicazione: (2025) -
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
di: Pahde, Frederik, et al.
Pubblicazione: (2025) -
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
di: Vielhaben, Johanna, et al.
Pubblicazione: (2024)