Linear Explanations for Individual Neurons
Fuente:
arXiv
Saved in:
| Main Authors: | Oikarinen, Tuomas, Weng, Tsui-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
by: Srivastava, Divyansh, et al.
Published: (2024)
by: Srivastava, Divyansh, et al.
Published: (2024)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
RAT: Boosting Misclassification Detection Ability without Extra Data
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Provably Robust Conformal Prediction with Improved Efficiency
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Crafting Large Language Models for Enhanced Interpretability
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Guaranteed Optimal Compositional Explanations for Neurons
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
Open Vocabulary Compositional Explanations for Neuron Alignment
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Concept Bottleneck Large Language Models
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Applying Graph Explanation to Operator Fusion
by: Mills, Keith G., et al.
Published: (2024)
by: Mills, Keith G., et al.
Published: (2024)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
by: Liao, Yi, et al.
Published: (2024)
by: Liao, Yi, et al.
Published: (2024)
Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations
by: Babadi, Narges, et al.
Published: (2026)
by: Babadi, Narges, et al.
Published: (2026)
DiG-IN: Diffusion Guidance for Investigating Networks -- Uncovering Classifier Differences Neuron Visualisations and Visual Counterfactual Explanations
by: Augustin, Maximilian, et al.
Published: (2023)
by: Augustin, Maximilian, et al.
Published: (2023)
Activation Matching for Explanation Generation
by: Suhail, Pirzada, et al.
Published: (2025)
by: Suhail, Pirzada, et al.
Published: (2025)
TACE: Tumor-Aware Counterfactual Explanations
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
The Manifold Hypothesis for Gradient-Based Explanations
by: Bordt, Sebastian, et al.
Published: (2022)
by: Bordt, Sebastian, et al.
Published: (2022)
Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape
by: Chen, Tiejin, et al.
Published: (2024)
by: Chen, Tiejin, et al.
Published: (2024)
What is Missing? Explaining Neurons Activated by Absent Concepts
by: Hesse, Robin, et al.
Published: (2026)
by: Hesse, Robin, et al.
Published: (2026)
Neuron Incidence Redistribution for Fairness in Medical Image Classification
by: Shoby, Abin, et al.
Published: (2026)
by: Shoby, Abin, et al.
Published: (2026)
On the Generalization and Causal Explanation in Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
How to Squeeze An Explanation Out of Your Model
by: Roxo, Tiago, et al.
Published: (2024)
by: Roxo, Tiago, et al.
Published: (2024)
Robustness of Visual Explanations to Common Data Augmentation
by: Tětková, Lenka, et al.
Published: (2023)
by: Tětková, Lenka, et al.
Published: (2023)
Neuronal Competition Groups with Supervised STDP for Spike-Based Classification
by: Goupy, Gaspard, et al.
Published: (2024)
by: Goupy, Gaspard, et al.
Published: (2024)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
NERO: Explainable Out-of-Distribution Detection with Neuron-level Relevance
by: Chhetri, Anju, et al.
Published: (2025)
by: Chhetri, Anju, et al.
Published: (2025)
Discovering and Mitigating Visual Biases through Keyword Explanation
by: Kim, Younghyun, et al.
Published: (2023)
by: Kim, Younghyun, et al.
Published: (2023)
Why Does It Look There? Structured Explanations for Image Classification
by: Li, Jiarui, et al.
Published: (2026)
by: Li, Jiarui, et al.
Published: (2026)
Architecture-Aware Explanation Auditing for Industrial Visual Inspection
by: Jia, Sibo, et al.
Published: (2026)
by: Jia, Sibo, et al.
Published: (2026)
xMIL: Insightful Explanations for Multiple Instance Learning in Histopathology
by: Hense, Julius, et al.
Published: (2024)
by: Hense, Julius, et al.
Published: (2024)
Data-centric Prediction Explanation via Kernelized Stein Discrepancy
by: Sarvmaili, Mahtab, et al.
Published: (2024)
by: Sarvmaili, Mahtab, et al.
Published: (2024)
VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
by: Srivastava, Aadi, et al.
Published: (2025)
by: Srivastava, Aadi, et al.
Published: (2025)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
On the Road with 16 Neurons: Mental Imagery with Bio-inspired Deep Neural Networks
by: Plebe, Alice, et al.
Published: (2020)
by: Plebe, Alice, et al.
Published: (2020)
Similar Items
-
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024) -
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026) -
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025) -
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025) -
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
by: Oikarinen, Tuomas, et al.
Published: (2025)