Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
Fuente:
arXiv
Saved in:
| Main Authors: | Pahde, Frederik, Wiegand, Thomas, Lapuschkin, Sebastian, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022)
by: Pahde, Frederik, et al.
Published: (2022)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)
by: Dawoud, Karam, et al.
Published: (2023)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
by: Tinauer, Christian, et al.
Published: (2024)
by: Tinauer, Christian, et al.
Published: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
by: Mekala, Rohan Reddy, et al.
Published: (2024)
by: Mekala, Rohan Reddy, et al.
Published: (2024)
Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection
by: Cantú, Emílio Dolgener, et al.
Published: (2025)
by: Cantú, Emílio Dolgener, et al.
Published: (2025)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
XAI-guided Insulator Anomaly Detection for Imbalanced Datasets
by: Hoefler, Maximilian Andreas, et al.
Published: (2024)
by: Hoefler, Maximilian Andreas, et al.
Published: (2024)
Playing the network backward: A Game Theoretic Attribution Framework
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
MIMM-X: Disentangling Spurious Correlations for Medical Image Analysis
by: Fay, Louisa, et al.
Published: (2025)
by: Fay, Louisa, et al.
Published: (2025)
Sparse, Efficient and Explainable Data Attribution with DualXDA
by: Yolcu, Galip Ümit, et al.
Published: (2024)
by: Yolcu, Galip Ümit, et al.
Published: (2024)
RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models
by: Varma, Maya, et al.
Published: (2024)
by: Varma, Maya, et al.
Published: (2024)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
by: Liu, Chenruo, et al.
Published: (2025)
by: Liu, Chenruo, et al.
Published: (2025)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
by: Ahmad, Kaleem
Published: (2025)
by: Ahmad, Kaleem
Published: (2025)
Mitigating Spurious Negative Pairs for Robust Industrial Anomaly Detection
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
by: Kang, Solha, et al.
Published: (2025)
by: Kang, Solha, et al.
Published: (2025)
Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces
by: Chen, Zhiling, et al.
Published: (2024)
by: Chen, Zhiling, et al.
Published: (2024)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Mechanistic understanding and validation of large AI models with SemanticLens
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
by: Xu, Juangui, et al.
Published: (2025)
by: Xu, Juangui, et al.
Published: (2025)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
SLIM: Spuriousness Mitigation with Minimal Human Annotations
by: Xuan, Xiwei, et al.
Published: (2024)
by: Xuan, Xiwei, et al.
Published: (2024)
Focusing Image Generation to Mitigate Spurious Correlations
by: Li, Xuewei, et al.
Published: (2024)
by: Li, Xuewei, et al.
Published: (2024)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
by: Jin, Qixuan, et al.
Published: (2024)
by: Jin, Qixuan, et al.
Published: (2024)
[Re] Improving Interpretation Faithfulness for Vision Transformers
by: Kurek, Izabela, et al.
Published: (2025)
by: Kurek, Izabela, et al.
Published: (2025)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Concept-based explanations of Segmentation and Detection models in Natural Disaster Management
by: Heydari, Samar, et al.
Published: (2026)
by: Heydari, Samar, et al.
Published: (2026)
Iterative Inference in a Chess-Playing Neural Network
by: Sandmann, Elias, et al.
Published: (2025)
by: Sandmann, Elias, et al.
Published: (2025)
Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis
by: Wang, Hongmei, et al.
Published: (2024)
by: Wang, Hongmei, et al.
Published: (2024)
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
by: Min, Hyungi, et al.
Published: (2026)
by: Min, Hyungi, et al.
Published: (2026)
Similar Items
-
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024) -
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025) -
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023) -
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)