Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Erogullari, Eren, Lapuschkin, Sebastian, Samek, Wojciech, Pahde, Frederik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025)
by: Pahde, Frederik, et al.
Published: (2025)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022)
by: Pahde, Frederik, et al.
Published: (2022)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)
by: Dawoud, Karam, et al.
Published: (2023)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
by: Tinauer, Christian, et al.
Published: (2024)
by: Tinauer, Christian, et al.
Published: (2024)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
On Background Bias of Post-Hoc Concept Embeddings in Computer Vision DNNs
by: Schwalbe, Gesina, et al.
Published: (2025)
by: Schwalbe, Gesina, et al.
Published: (2025)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
by: Kim, Hyeonjin, et al.
Published: (2026)
by: Kim, Hyeonjin, et al.
Published: (2026)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
OmniPrism: Learning Disentangled Visual Concept for Image Generation
by: Li, Yangyang, et al.
Published: (2024)
by: Li, Yangyang, et al.
Published: (2024)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
by: Liu, Chenruo, et al.
Published: (2025)
by: Liu, Chenruo, et al.
Published: (2025)
Playing the network backward: A Game Theoretic Attribution Framework
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability
by: Mikriukov, Georgii, et al.
Published: (2023)
by: Mikriukov, Georgii, et al.
Published: (2023)
A Geometric Unification of Concept Learning with Concept Cones
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
by: Petsiuk, Vitali, et al.
Published: (2024)
by: Petsiuk, Vitali, et al.
Published: (2024)
Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
by: Chang, Wonjoon, et al.
Published: (2023)
by: Chang, Wonjoon, et al.
Published: (2023)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
by: Achtibat, Reduan, et al.
Published: (2022)
by: Achtibat, Reduan, et al.
Published: (2022)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
by: Kwon, Gihyun, et al.
Published: (2024)
by: Kwon, Gihyun, et al.
Published: (2024)
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
by: Lin, Shen, et al.
Published: (2026)
by: Lin, Shen, et al.
Published: (2026)
Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
by: Pawlonka, Szymon, et al.
Published: (2025)
by: Pawlonka, Szymon, et al.
Published: (2025)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
by: Mekala, Rohan Reddy, et al.
Published: (2024)
by: Mekala, Rohan Reddy, et al.
Published: (2024)
Feature Attribution Stability Suite: How Stable Are Post-Hoc Attributions?
by: Subramaniakuppusamy, Kamalasankari, et al.
Published: (2026)
by: Subramaniakuppusamy, Kamalasankari, et al.
Published: (2026)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
by: Rao, Sukrut, et al.
Published: (2024)
by: Rao, Sukrut, et al.
Published: (2024)
ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
by: Chavhan, Ruchika, et al.
Published: (2024)
by: Chavhan, Ruchika, et al.
Published: (2024)
Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification
by: De Santis, Antonio, et al.
Published: (2024)
by: De Santis, Antonio, et al.
Published: (2024)
Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretations
by: Xu, Xinyue, et al.
Published: (2024)
by: Xu, Xinyue, et al.
Published: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
by: Elisha, Yehonatan, et al.
Published: (2026)
by: Elisha, Yehonatan, et al.
Published: (2026)
Editable Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
by: Caselli, Lorenzo, et al.
Published: (2026)
by: Caselli, Lorenzo, et al.
Published: (2026)
Symbolic Disentangled Representations for Images
by: Korchemnyi, Alexandr, et al.
Published: (2024)
by: Korchemnyi, Alexandr, et al.
Published: (2024)
Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware Optimization
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Similar Items
-
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025) -
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024) -
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023) -
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)