Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
Fuente:
arXiv
Saved in:
| Main Authors: | Oikarinen, Tuomas, Yan, Ge, Weng, Tsui-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024)
by: Oikarinen, Tuomas, et al.
Published: (2024)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
Crafting Large Language Models for Enhanced Interpretability
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
Concept Bottleneck Large Language Models
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Sanity Checks for Agentic Data Science
by: Rewolinski, Zachary T., et al.
Published: (2026)
by: Rewolinski, Zachary T., et al.
Published: (2026)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
RAT: Boosting Misclassification Detection Ability without Extra Data
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
by: Srivastava, Divyansh, et al.
Published: (2024)
by: Srivastava, Divyansh, et al.
Published: (2024)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
by: Sun, Chung-En, et al.
Published: (2025)
by: Sun, Chung-En, et al.
Published: (2025)
Provably Robust Conformal Prediction with Improved Efficiency
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Sanity Checking Causal Representation Learning on a Simple Real-World System
by: Gamella, Juan L., et al.
Published: (2025)
by: Gamella, Juan L., et al.
Published: (2025)
Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Towards a Unified Framework for Evaluating Explanations
by: Pinto, Juan D., et al.
Published: (2024)
by: Pinto, Juan D., et al.
Published: (2024)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
CoSy: Evaluating Textual Explanations of Neurons
by: Kopf, Laura, et al.
Published: (2024)
by: Kopf, Laura, et al.
Published: (2024)
Distance Marching for Generative Modeling
by: Wang, Zimo, et al.
Published: (2026)
by: Wang, Zimo, et al.
Published: (2026)
Effective Skill Unlearning through Intervention and Abstention
by: Li, Yongce, et al.
Published: (2025)
by: Li, Yongce, et al.
Published: (2025)
Breaking the Barrier: Enhanced Utility and Robustness in Smoothed DRL Agents
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Graph Concept Bottleneck Models
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Prediction without Preclusion: Recourse Verification with Reachable Sets
by: Kothari, Avni, et al.
Published: (2023)
by: Kothari, Avni, et al.
Published: (2023)
Abstracted Shapes as Tokens -- A Generalizable and Interpretable Model for Time-series Classification
by: Wen, Yunshi, et al.
Published: (2024)
by: Wen, Yunshi, et al.
Published: (2024)
Statistical Inference for Responsiveness Verification
by: Cheon, Seung Hyun, et al.
Published: (2025)
by: Cheon, Seung Hyun, et al.
Published: (2025)
Understanding Fixed Predictions via Confined Regions
by: Lawless, Connor, et al.
Published: (2025)
by: Lawless, Connor, et al.
Published: (2025)
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
by: Zhao, Haiyan, et al.
Published: (2026)
by: Zhao, Haiyan, et al.
Published: (2026)
A Unified Evaluation Framework for Epistemic Predictions
by: Manchingal, Shireen Kudukkil, et al.
Published: (2025)
by: Manchingal, Shireen Kudukkil, et al.
Published: (2025)
Guaranteed Optimal Compositional Explanations for Neurons
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
Beyond Attribution: Unified Concept-Level Explanations
by: Liu, Junhao, et al.
Published: (2024)
by: Liu, Junhao, et al.
Published: (2024)
Unified Explanations in Machine Learning Models: A Perturbation Approach
by: Dineen, Jacob, et al.
Published: (2024)
by: Dineen, Jacob, et al.
Published: (2024)
Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced Data
by: Weng, Pei-Yau, et al.
Published: (2025)
by: Weng, Pei-Yau, et al.
Published: (2025)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
by: Guo, Yuhan, et al.
Published: (2025)
by: Guo, Yuhan, et al.
Published: (2025)
Open Vocabulary Compositional Explanations for Neuron Alignment
by: La Rosa, Biagio, et al.
Published: (2025)
by: La Rosa, Biagio, et al.
Published: (2025)
A Sanity Check for AI-generated Image Detection
by: Yan, Shilin, et al.
Published: (2024)
by: Yan, Shilin, et al.
Published: (2024)
Similar Items
-
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025) -
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024) -
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025) -
Crafting Large Language Models for Enhanced Interpretability
by: Sun, Chung-En, et al.
Published: (2024) -
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)