Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
Fuente:
arXiv
Guardado en:
| Autores principales: | Pahde, Frederik, Dreyer, Maximilian, Weber, Leander, Weckbecker, Moritz, Anders, Christopher J., Wiegand, Thomas, Samek, Wojciech, Lapuschkin, Sebastian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
por: Pahde, Frederik, et al.
Publicado: (2025)
por: Pahde, Frederik, et al.
Publicado: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
por: Erogullari, Eren, et al.
Publicado: (2025)
por: Erogullari, Eren, et al.
Publicado: (2025)
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
por: Weber, Leander, et al.
Publicado: (2023)
por: Weber, Leander, et al.
Publicado: (2023)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Sparse, Efficient and Explainable Data Attribution with DualXDA
por: Yolcu, Galip Ümit, et al.
Publicado: (2024)
por: Yolcu, Galip Ümit, et al.
Publicado: (2024)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
por: Dreyer, Maximilian, et al.
Publicado: (2023)
por: Dreyer, Maximilian, et al.
Publicado: (2023)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
por: Achtibat, Reduan, et al.
Publicado: (2022)
por: Achtibat, Reduan, et al.
Publicado: (2022)
PINNfluence: Influence Functions for Physics-Informed Neural Networks
por: Naujoks, Jonas R., et al.
Publicado: (2024)
por: Naujoks, Jonas R., et al.
Publicado: (2024)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
por: Dreyer, Maximilian, et al.
Publicado: (2025)
por: Dreyer, Maximilian, et al.
Publicado: (2025)
Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks
por: Naujoks, Jonas R., et al.
Publicado: (2025)
por: Naujoks, Jonas R., et al.
Publicado: (2025)
Mechanistic understanding and validation of large AI models with SemanticLens
por: Dreyer, Maximilian, et al.
Publicado: (2025)
por: Dreyer, Maximilian, et al.
Publicado: (2025)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
por: Gururaj, Shreyas, et al.
Publicado: (2025)
por: Gururaj, Shreyas, et al.
Publicado: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
por: Dreyer, Maximilian, et al.
Publicado: (2024)
por: Dreyer, Maximilian, et al.
Publicado: (2024)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
por: Bouanani, Oussama, et al.
Publicado: (2026)
por: Bouanani, Oussama, et al.
Publicado: (2026)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
por: Becking, Daniel, et al.
Publicado: (2021)
por: Becking, Daniel, et al.
Publicado: (2021)
Structural Compactness as a Complementary Criterion for Explanation Quality
por: Mesgari, Mohammad Mahdi, et al.
Publicado: (2026)
por: Mesgari, Mohammad Mahdi, et al.
Publicado: (2026)
Iterative Inference in a Chess-Playing Neural Network
por: Sandmann, Elias, et al.
Publicado: (2025)
por: Sandmann, Elias, et al.
Publicado: (2025)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
por: Achtibat, Reduan, et al.
Publicado: (2024)
por: Achtibat, Reduan, et al.
Publicado: (2024)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
See What I Mean? CUE: A Cognitive Model of Understanding Explanations
por: Labarta, Tobias, et al.
Publicado: (2025)
por: Labarta, Tobias, et al.
Publicado: (2025)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
por: Tinauer, Christian, et al.
Publicado: (2024)
por: Tinauer, Christian, et al.
Publicado: (2024)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025)
por: Hufe, Lorenz, et al.
Publicado: (2025)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
Attribution-Guided Decoding
por: Komorowski, Piotr, et al.
Publicado: (2025)
por: Komorowski, Piotr, et al.
Publicado: (2025)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
por: Hedström, Anna, et al.
Publicado: (2024)
por: Hedström, Anna, et al.
Publicado: (2024)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
por: Bley, Florian, et al.
Publicado: (2024)
por: Bley, Florian, et al.
Publicado: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
Circuit Insights: Towards Interpretability Beyond Activations
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
FADE: Why Bad Descriptions Happen to Good Features
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
X-SYS: A Reference Architecture for Interactive Explanation Systems
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
$α$-TCAV: A Unified Framework for Testing with Concept Activation Vectors
por: Schnoor, Ekkehard, et al.
Publicado: (2026)
por: Schnoor, Ekkehard, et al.
Publicado: (2026)
LieSolver: A PDE-constrained solver for IBVPs using Lie symmetries
por: Klausen, René P., et al.
Publicado: (2025)
por: Klausen, René P., et al.
Publicado: (2025)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Human-Centered Evaluation of XAI Methods
por: Dawoud, Karam, et al.
Publicado: (2023)
por: Dawoud, Karam, et al.
Publicado: (2023)
A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
por: Arras, Leila, et al.
Publicado: (2025)
por: Arras, Leila, et al.
Publicado: (2025)
A Fresh Look at Sanity Checks for Saliency Maps
por: Hedström, Anna, et al.
Publicado: (2024)
por: Hedström, Anna, et al.
Publicado: (2024)
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
por: Ostermann, Simon, et al.
Publicado: (2026)
por: Ostermann, Simon, et al.
Publicado: (2026)
Playing the network backward: A Game Theoretic Attribution Framework
por: Zimmermann, Jakob Paul, et al.
Publicado: (2026)
por: Zimmermann, Jakob Paul, et al.
Publicado: (2026)
Ejemplares similares
-
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
por: Pahde, Frederik, et al.
Publicado: (2025) -
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
por: Erogullari, Eren, et al.
Publicado: (2025) -
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
por: Weber, Leander, et al.
Publicado: (2023) -
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024) -
Sparse, Efficient and Explainable Data Attribution with DualXDA
por: Yolcu, Galip Ümit, et al.
Publicado: (2024)