Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
Fuente:
arXiv
Salvato in:
| Autori principali: | Bareeva, Dilyara, Dreyer, Maximilian, Pahde, Frederik, Samek, Wojciech, Lapuschkin, Sebastian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
di: Pahde, Frederik, et al.
Pubblicazione: (2025)
di: Pahde, Frederik, et al.
Pubblicazione: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
di: Dreyer, Maximilian, et al.
Pubblicazione: (2023)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2023)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
di: Pahde, Frederik, et al.
Pubblicazione: (2022)
di: Pahde, Frederik, et al.
Pubblicazione: (2022)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
di: Vielhaben, Johanna, et al.
Pubblicazione: (2024)
di: Vielhaben, Johanna, et al.
Pubblicazione: (2024)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
di: Bouanani, Oussama, et al.
Pubblicazione: (2026)
di: Bouanani, Oussama, et al.
Pubblicazione: (2026)
Towards Visually Explaining Statistical Tests with Applications in Biomedical Imaging
di: Javanbakhat, Masoumeh, et al.
Pubblicazione: (2026)
di: Javanbakhat, Masoumeh, et al.
Pubblicazione: (2026)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
di: Tinauer, Christian, et al.
Pubblicazione: (2024)
di: Tinauer, Christian, et al.
Pubblicazione: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
Manipulating Feature Visualizations with Gradient Slingshots
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2024)
di: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Pubblicazione: (2024)
Human-Centered Evaluation of XAI Methods
di: Dawoud, Karam, et al.
Pubblicazione: (2023)
di: Dawoud, Karam, et al.
Pubblicazione: (2023)
Playing the network backward: A Game Theoretic Attribution Framework
di: Zimmermann, Jakob Paul, et al.
Pubblicazione: (2026)
di: Zimmermann, Jakob Paul, et al.
Pubblicazione: (2026)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
di: Mekala, Rohan Reddy, et al.
Pubblicazione: (2024)
di: Mekala, Rohan Reddy, et al.
Pubblicazione: (2024)
XAI-guided Insulator Anomaly Detection for Imbalanced Datasets
di: Hoefler, Maximilian Andreas, et al.
Pubblicazione: (2024)
di: Hoefler, Maximilian Andreas, et al.
Pubblicazione: (2024)
MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
di: Kang, Changwon, et al.
Pubblicazione: (2025)
di: Kang, Changwon, et al.
Pubblicazione: (2025)
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
di: Bareeva, Dilyara, et al.
Pubblicazione: (2024)
Model Guidance via Explanations Turns Image Classifiers into Segmentation Models
di: Yu, Xiaoyan, et al.
Pubblicazione: (2024)
di: Yu, Xiaoyan, et al.
Pubblicazione: (2024)
Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection
di: Cantú, Emílio Dolgener, et al.
Pubblicazione: (2025)
di: Cantú, Emílio Dolgener, et al.
Pubblicazione: (2025)
The Bias of Harmful Label Associations in Vision-Language Models
di: Hazirbas, Caner, et al.
Pubblicazione: (2024)
di: Hazirbas, Caner, et al.
Pubblicazione: (2024)
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
di: Tan, Xuwei, et al.
Pubblicazione: (2026)
di: Tan, Xuwei, et al.
Pubblicazione: (2026)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
di: Cho, Yooshin, et al.
Pubblicazione: (2025)
di: Cho, Yooshin, et al.
Pubblicazione: (2025)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
di: Becking, Daniel, et al.
Pubblicazione: (2021)
di: Becking, Daniel, et al.
Pubblicazione: (2021)
From Attribution to Action: A Human-Centered Application of Activation Steering
di: Labarta, Tobias, et al.
Pubblicazione: (2026)
di: Labarta, Tobias, et al.
Pubblicazione: (2026)
Low-Rank Adaptation with Task-Relevant Feature Enhancement for Fine-tuning Language Models
di: Li, Changqun, et al.
Pubblicazione: (2024)
di: Li, Changqun, et al.
Pubblicazione: (2024)
Multi Attribute Bias Mitigation via Representation Learning
di: Dwivedi, Rajeev Ranjan, et al.
Pubblicazione: (2025)
di: Dwivedi, Rajeev Ranjan, et al.
Pubblicazione: (2025)
BAdd: Bias Mitigation through Bias Addition
di: Sarridis, Ioannis, et al.
Pubblicazione: (2024)
di: Sarridis, Ioannis, et al.
Pubblicazione: (2024)
Mitigating Low-Frequency Bias: Feature Recalibration and Frequency Attention Regularization for Adversarial Robustness
di: Zhang, Kejia, et al.
Pubblicazione: (2024)
di: Zhang, Kejia, et al.
Pubblicazione: (2024)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
di: Liu, Yule, et al.
Pubblicazione: (2026)
di: Liu, Yule, et al.
Pubblicazione: (2026)
Mitigating Bias Using Model-Agnostic Data Attribution
di: De Coninck, Sander, et al.
Pubblicazione: (2024)
di: De Coninck, Sander, et al.
Pubblicazione: (2024)
Frequency Regulation for Exposure Bias Mitigation in Diffusion Models
di: Yu, Meng, et al.
Pubblicazione: (2025)
di: Yu, Meng, et al.
Pubblicazione: (2025)
Common-Sense Bias Modeling for Classification Tasks
di: Zhang, Miao, et al.
Pubblicazione: (2024)
di: Zhang, Miao, et al.
Pubblicazione: (2024)
Fractional Diffusion Bridge Models
di: Nobis, Gabriel, et al.
Pubblicazione: (2025)
di: Nobis, Gabriel, et al.
Pubblicazione: (2025)
ViG-Bias: Visually Grounded Bias Discovery and Mitigation
di: Marani, Badr-Eddine, et al.
Pubblicazione: (2024)
di: Marani, Badr-Eddine, et al.
Pubblicazione: (2024)
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
di: Tian, Bowei, et al.
Pubblicazione: (2024)
di: Tian, Bowei, et al.
Pubblicazione: (2024)
Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation
di: Wijenayake, Buddhi, et al.
Pubblicazione: (2026)
di: Wijenayake, Buddhi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
di: Pahde, Frederik, et al.
Pubblicazione: (2025) -
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
di: Erogullari, Eren, et al.
Pubblicazione: (2025) -
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
di: Dreyer, Maximilian, et al.
Pubblicazione: (2023) -
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
di: Pahde, Frederik, et al.
Pubblicazione: (2022)