Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Kulkarni, Akshay, Weng, Tsui-Wei, Narayanaswamy, Vivek, Liu, Shusen, Sakla, Wesam A., Thopalli, Kowshik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
On the Use of Anchoring for Training Vision Models
by: Narayanaswamy, Vivek, et al.
Published: (2024)
by: Narayanaswamy, Vivek, et al.
Published: (2024)
Leveraging Registers in Vision Transformers for Robust Adaptation
by: Yellapragada, Srikar, et al.
Published: (2025)
by: Yellapragada, Srikar, et al.
Published: (2025)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
by: Flora, James, et al.
Published: (2026)
by: Flora, James, et al.
Published: (2026)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
by: Srivastava, Divyansh, et al.
Published: (2024)
by: Srivastava, Divyansh, et al.
Published: (2024)
DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
by: Subramanyam, Rakshith, et al.
Published: (2024)
by: Subramanyam, Rakshith, et al.
Published: (2024)
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
Speeding Up Image Classifiers with Little Companions
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024)
by: Oikarinen, Tuomas, et al.
Published: (2024)
Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification
by: Tong, Schrasing, et al.
Published: (2026)
by: Tong, Schrasing, et al.
Published: (2026)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
by: Wesp, Philipp, et al.
Published: (2026)
by: Wesp, Philipp, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Concept Bottleneck Models Without Predefined Concepts
by: Schrodi, Simon, et al.
Published: (2024)
by: Schrodi, Simon, et al.
Published: (2024)
RAT: Boosting Misclassification Detection Ability without Extra Data
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Hyperbolic Concept Bottleneck Models
by: Uyterlinde, Daniel, et al.
Published: (2026)
by: Uyterlinde, Daniel, et al.
Published: (2026)
Flexible Concept Bottleneck Model
by: Du, Xingbo, et al.
Published: (2025)
by: Du, Xingbo, et al.
Published: (2025)
Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations
by: Xu, Xinyue, et al.
Published: (2024)
by: Xu, Xinyue, et al.
Published: (2024)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
by: Yeung, Calvin, et al.
Published: (2026)
by: Yeung, Calvin, et al.
Published: (2026)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Towards Faithful Multimodal Concept Bottleneck Models
by: Moreau, Pierre, et al.
Published: (2026)
by: Moreau, Pierre, et al.
Published: (2026)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Provably Robust Conformal Prediction with Improved Efficiency
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Debugging Concept Bottleneck Models through Removal and Retraining
by: Enouen, Eric, et al.
Published: (2025)
by: Enouen, Eric, et al.
Published: (2025)
Editable Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
by: Liu, Christina, et al.
Published: (2025)
by: Liu, Christina, et al.
Published: (2025)
Learning a Clinically-Relevant Concept Bottleneck for Lesion Detection in Breast Ultrasound
by: Bunnell, Arianna, et al.
Published: (2024)
by: Bunnell, Arianna, et al.
Published: (2024)
Zero-shot Concept Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Process-Guided Concept Bottleneck Model
by: Asiyabi, Reza M., et al.
Published: (2026)
by: Asiyabi, Reza M., et al.
Published: (2026)
Semi-supervised Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
by: Huang, Gaoxiang, et al.
Published: (2025)
by: Huang, Gaoxiang, et al.
Published: (2025)
Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset
by: Nápoles, Gonzalo, et al.
Published: (2026)
by: Nápoles, Gonzalo, et al.
Published: (2026)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Similar Items
-
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026) -
On the Use of Anchoring for Training Vision Models
by: Narayanaswamy, Vivek, et al.
Published: (2024) -
Leveraging Registers in Vision Transformers for Robust Adaptation
by: Yellapragada, Srikar, et al.
Published: (2025) -
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
by: Flora, James, et al.
Published: (2026) -
Interpretable Generative Models through Post-hoc Concept Bottlenecks
by: Kulkarni, Akshay, et al.
Published: (2025)