Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Bağcı, Doğukan, Schiele, Bernt, Schaub-Meyer, Simone, Fischer, Jonas, Hesse, Robin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangling Polysemantic Channels in Convolutional Neural Networks
by: Hesse, Robin, et al.
Published: (2025)
by: Hesse, Robin, et al.
Published: (2025)
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
by: Hesse, Robin, et al.
Published: (2025)
by: Hesse, Robin, et al.
Published: (2025)
What is Missing? Explaining Neurons Activated by Absent Concepts
by: Hesse, Robin, et al.
Published: (2026)
by: Hesse, Robin, et al.
Published: (2026)
Benchmarking the Attribution Quality of Vision Models
by: Hesse, Robin, et al.
Published: (2024)
by: Hesse, Robin, et al.
Published: (2024)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
Disentangling Polysemantic Neurons with a Null-Calibrated Polysemanticity Index and Causal Patch Interventions
by: Gupta, Manan, et al.
Published: (2025)
by: Gupta, Manan, et al.
Published: (2025)
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
by: Parchami-Araghi, Amin, et al.
Published: (2025)
by: Parchami-Araghi, Amin, et al.
Published: (2025)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
by: Arya, Shreyash, et al.
Published: (2024)
by: Arya, Shreyash, et al.
Published: (2024)
CFM: Language-aligned Concept Foundation Model for Vision
by: Wittenmayer, Kai, et al.
Published: (2026)
by: Wittenmayer, Kai, et al.
Published: (2026)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
Adaptive Hierarchical Certification for Segmentation using Randomized Smoothing
by: Anani, Alaa, et al.
Published: (2024)
by: Anani, Alaa, et al.
Published: (2024)
GLASS: Guided Latent Slot Diffusion for Object-Centric Learning
by: Singh, Krishnakant, et al.
Published: (2024)
by: Singh, Krishnakant, et al.
Published: (2024)
Evaluating Object-Centric Models beyond Object Discovery
by: Singh, Krishnakant, et al.
Published: (2026)
by: Singh, Krishnakant, et al.
Published: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Adversarial Training against Location-Optimized Adversarial Patches
by: Rao, Sukrut, et al.
Published: (2020)
by: Rao, Sukrut, et al.
Published: (2020)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
by: Pham, Nhi, et al.
Published: (2025)
by: Pham, Nhi, et al.
Published: (2025)
Spatial Reasoners for Continuous Variables in Any Domain
by: Pogodzinski, Bart, et al.
Published: (2025)
by: Pogodzinski, Bart, et al.
Published: (2025)
Spatial Reasoning with Denoising Models
by: Wewer, Christopher, et al.
Published: (2025)
by: Wewer, Christopher, et al.
Published: (2025)
Test-Time Visual In-Context Tuning
by: Xie, Jiahao, et al.
Published: (2025)
by: Xie, Jiahao, et al.
Published: (2025)
Pixel-level Certified Explanations via Randomized Smoothing
by: Anani, Alaa, et al.
Published: (2025)
by: Anani, Alaa, et al.
Published: (2025)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
by: Rao, Sukrut, et al.
Published: (2024)
by: Rao, Sukrut, et al.
Published: (2024)
Activation Subspaces for Out-of-Distribution Detection
by: Zöngür, Barış, et al.
Published: (2025)
by: Zöngür, Barış, et al.
Published: (2025)
Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
by: Wille, Melanie, et al.
Published: (2025)
by: Wille, Melanie, et al.
Published: (2025)
Studying How to Efficiently and Effectively Guide Models with Explanations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
by: Zhao, Yue, et al.
Published: (2022)
by: Zhao, Yue, et al.
Published: (2022)
AnyUp: Universal Feature Upsampling
by: Wimmer, Thomas, et al.
Published: (2025)
by: Wimmer, Thomas, et al.
Published: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
by: Alshami, Eyad, et al.
Published: (2025)
by: Alshami, Eyad, et al.
Published: (2025)
MEt3R: Measuring Multi-View Consistency in Generated Images
by: Asim, Mohammad, et al.
Published: (2025)
by: Asim, Mohammad, et al.
Published: (2025)
On the Robustness Tradeoff in Fine-Tuning
by: Li, Kunyang, et al.
Published: (2025)
by: Li, Kunyang, et al.
Published: (2025)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
by: Kuzucu, Selim, et al.
Published: (2026)
by: Kuzucu, Selim, et al.
Published: (2026)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
by: Anani, Alaa, et al.
Published: (2026)
by: Anani, Alaa, et al.
Published: (2026)
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
by: Wróbel, Adam, et al.
Published: (2026)
by: Wróbel, Adam, et al.
Published: (2026)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
by: Chen, Hao, et al.
Published: (2022)
by: Chen, Hao, et al.
Published: (2022)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
by: Shi, Shaoshuai, et al.
Published: (2023)
by: Shi, Shaoshuai, et al.
Published: (2023)
Improved Segmentation of Polyps and Visual Explainability Analysis
by: Asare, Akwasi, et al.
Published: (2025)
by: Asare, Akwasi, et al.
Published: (2025)
Similar Items
-
Disentangling Polysemantic Channels in Convolutional Neural Networks
by: Hesse, Robin, et al.
Published: (2025) -
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
by: Hesse, Robin, et al.
Published: (2025) -
What is Missing? Explaining Neurons Activated by Absent Concepts
by: Hesse, Robin, et al.
Published: (2026) -
Benchmarking the Attribution Quality of Vision Models
by: Hesse, Robin, et al.
Published: (2024) -
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)