Mechanistic understanding and validation of large AI models with SemanticLens
Fuente:
arXiv
Saved in:
| Main Authors: | Dreyer, Maximilian, Berend, Jim, Labarta, Tobias, Vielhaben, Johanna, Wiegand, Thomas, Lapuschkin, Sebastian, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
by: Labarta, Tobias, et al.
Published: (2026)
by: Labarta, Tobias, et al.
Published: (2026)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
X-SYS: A Reference Architecture for Interactive Explanation Systems
by: Labarta, Tobias, et al.
Published: (2026)
by: Labarta, Tobias, et al.
Published: (2026)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025)
by: Pahde, Frederik, et al.
Published: (2025)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
by: Achtibat, Reduan, et al.
Published: (2022)
by: Achtibat, Reduan, et al.
Published: (2022)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
by: Weber, Leander, et al.
Published: (2023)
by: Weber, Leander, et al.
Published: (2023)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Sparse, Efficient and Explainable Data Attribution with DualXDA
by: Yolcu, Galip Ümit, et al.
Published: (2024)
by: Yolcu, Galip Ümit, et al.
Published: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022)
by: Pahde, Frederik, et al.
Published: (2022)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2025)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2025)
Iterative Inference in a Chess-Playing Neural Network
by: Sandmann, Elias, et al.
Published: (2025)
by: Sandmann, Elias, et al.
Published: (2025)
FADE: Why Bad Descriptions Happen to Good Features
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
by: Bley, Florian, et al.
Published: (2024)
by: Bley, Florian, et al.
Published: (2024)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
See What I Mean? CUE: A Cognitive Model of Understanding Explanations
by: Labarta, Tobias, et al.
Published: (2025)
by: Labarta, Tobias, et al.
Published: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
LieSolver: A PDE-constrained solver for IBVPs using Lie symmetries
by: Klausen, René P., et al.
Published: (2025)
by: Klausen, René P., et al.
Published: (2025)
Structural Compactness as a Complementary Criterion for Explanation Quality
by: Mesgari, Mohammad Mahdi, et al.
Published: (2026)
by: Mesgari, Mohammad Mahdi, et al.
Published: (2026)
PINNfluence: Influence Functions for Physics-Informed Neural Networks
by: Naujoks, Jonas R., et al.
Published: (2024)
by: Naujoks, Jonas R., et al.
Published: (2024)
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)
by: Dawoud, Karam, et al.
Published: (2023)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
by: Gururaj, Shreyas, et al.
Published: (2025)
by: Gururaj, Shreyas, et al.
Published: (2025)
Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks
by: Naujoks, Jonas R., et al.
Published: (2025)
by: Naujoks, Jonas R., et al.
Published: (2025)
Circuit Insights: Towards Interpretability Beyond Activations
by: Golimblevskaia, Elena, et al.
Published: (2025)
by: Golimblevskaia, Elena, et al.
Published: (2025)
Model Science: getting serious about verification, explanation and control of AI systems
by: Biecek, Przemyslaw, et al.
Published: (2025)
by: Biecek, Przemyslaw, et al.
Published: (2025)
FeatInv: Spatially resolved mapping from feature space to input space using conditional diffusion models
by: Neukirch, Nils, et al.
Published: (2025)
by: Neukirch, Nils, et al.
Published: (2025)
Building Trust in PINNs: Error Estimation through Finite Difference Methods
by: Krasowski, Aleksander, et al.
Published: (2026)
by: Krasowski, Aleksander, et al.
Published: (2026)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
Playing the network backward: A Game Theoretic Attribution Framework
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
by: Tinauer, Christian, et al.
Published: (2024)
by: Tinauer, Christian, et al.
Published: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Decoupling Pixel Flipping and Occlusion Strategy for Consistent XAI Benchmarks
by: Blücher, Stefan, et al.
Published: (2024)
by: Blücher, Stefan, et al.
Published: (2024)
Similar Items
-
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025) -
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024) -
From Attribution to Action: A Human-Centered Application of Activation Steering
by: Labarta, Tobias, et al.
Published: (2026) -
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024) -
X-SYS: A Reference Architecture for Interactive Explanation Systems
by: Labarta, Tobias, et al.
Published: (2026)