Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhalla, Usha, Srinivas, Suraj, Lakkaraju, Himabindu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
di: Srinivas, Suraj, et al.
Pubblicazione: (2023)
di: Srinivas, Suraj, et al.
Pubblicazione: (2023)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Feature Attribution Stability Suite: How Stable Are Post-Hoc Attributions?
di: Subramaniakuppusamy, Kamalasankari, et al.
Pubblicazione: (2026)
di: Subramaniakuppusamy, Kamalasankari, et al.
Pubblicazione: (2026)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
Comprehensive Attribution: Inherently Explainable Vision Model with Feature Detector
di: Zhang, Xianren, et al.
Pubblicazione: (2024)
di: Zhang, Xianren, et al.
Pubblicazione: (2024)
PostHoc FREE Calibrating on Kolmogorov Arnold Networks
di: Liang, Wenhao, et al.
Pubblicazione: (2025)
di: Liang, Wenhao, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
di: Hsu, Chia-Hong, et al.
Pubblicazione: (2026)
di: Hsu, Chia-Hong, et al.
Pubblicazione: (2026)
Characterizing Data Point Vulnerability via Average-Case Robustness
di: Han, Tessa, et al.
Pubblicazione: (2023)
di: Han, Tessa, et al.
Pubblicazione: (2023)
Post-Hoc MOTS: Exploring the Capabilities of Time-Symmetric Multi-Object Tracking
di: Szabó, Gergely, et al.
Pubblicazione: (2024)
di: Szabó, Gergely, et al.
Pubblicazione: (2024)
Train Once, Forget Precisely: Anchored Optimization for Efficient Post-Hoc Unlearning
di: Sanga, Prabhav, et al.
Pubblicazione: (2025)
di: Sanga, Prabhav, et al.
Pubblicazione: (2025)
X-SiT: Inherently Interpretable Surface Vision Transformers for Dementia Diagnosis
di: Bongratz, Fabian, et al.
Pubblicazione: (2025)
di: Bongratz, Fabian, et al.
Pubblicazione: (2025)
Hessian Surgery: Class-Targeted Post-Hoc Rebalancing via Hessian Spike Perturbation
di: Vigna, Hugo, et al.
Pubblicazione: (2026)
di: Vigna, Hugo, et al.
Pubblicazione: (2026)
Interpretable Network Visualizations: A Human-in-the-Loop Approach for Post-hoc Explainability of CNN-based Image Classification
di: Bianchi, Matteo, et al.
Pubblicazione: (2024)
di: Bianchi, Matteo, et al.
Pubblicazione: (2024)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
di: Arya, Shreyash, et al.
Pubblicazione: (2024)
di: Arya, Shreyash, et al.
Pubblicazione: (2024)
Advancing Ante-Hoc Explainable Models through Generative Adversarial Networks
di: Garg, Tanmay, et al.
Pubblicazione: (2024)
di: Garg, Tanmay, et al.
Pubblicazione: (2024)
Generalized Group Data Attribution
di: Ley, Dan, et al.
Pubblicazione: (2024)
di: Ley, Dan, et al.
Pubblicazione: (2024)
Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification
di: De Santis, Antonio, et al.
Pubblicazione: (2024)
di: De Santis, Antonio, et al.
Pubblicazione: (2024)
Training Feature Attribution for Vision Models
di: Bacha, Aziz, et al.
Pubblicazione: (2025)
di: Bacha, Aziz, et al.
Pubblicazione: (2025)
Deep Image Segmentation via Discriminant Feature Learning
di: Sztamborski, Adam Dawid, et al.
Pubblicazione: (2026)
di: Sztamborski, Adam Dawid, et al.
Pubblicazione: (2026)
Attri-Net: A Globally and Locally Inherently Interpretable Model for Multi-Label Classification Using Class-Specific Counterfactuals
di: Sun, Susu, et al.
Pubblicazione: (2024)
di: Sun, Susu, et al.
Pubblicazione: (2024)
Interpretable Few-shot Learning with Online Attribute Selection
di: Zarei, Mohammad Reza, et al.
Pubblicazione: (2022)
di: Zarei, Mohammad Reza, et al.
Pubblicazione: (2022)
Evaluating Feature Attribution Methods in the Image Domain
di: Gevaert, Arne, et al.
Pubblicazione: (2022)
di: Gevaert, Arne, et al.
Pubblicazione: (2022)
The Gaussian Discriminant Variational Autoencoder (GdVAE): A Self-Explainable Model with Counterfactual Explanations
di: Haselhoff, Anselm, et al.
Pubblicazione: (2024)
di: Haselhoff, Anselm, et al.
Pubblicazione: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
di: Erogullari, Eren, et al.
Pubblicazione: (2025)
On Background Bias of Post-Hoc Concept Embeddings in Computer Vision DNNs
di: Schwalbe, Gesina, et al.
Pubblicazione: (2025)
di: Schwalbe, Gesina, et al.
Pubblicazione: (2025)
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
di: Chun, Sanghyuk
Pubblicazione: (2025)
di: Chun, Sanghyuk
Pubblicazione: (2025)
Fast Post-Hoc Confidence Fusion for 3-Class Open-Set Aerial Object Detection
di: Loukovitis, Spyridon, et al.
Pubblicazione: (2025)
di: Loukovitis, Spyridon, et al.
Pubblicazione: (2025)
Self-Supervised Discriminative Feature Learning for Deep Multi-View Clustering
di: Xu, Jie, et al.
Pubblicazione: (2021)
di: Xu, Jie, et al.
Pubblicazione: (2021)
A Generic Self-Supervised Framework of Learning Invariant Discriminative Features
di: Ntelemis, Foivos, et al.
Pubblicazione: (2022)
di: Ntelemis, Foivos, et al.
Pubblicazione: (2022)
Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation
di: Liu, Fengbei, et al.
Pubblicazione: (2023)
di: Liu, Fengbei, et al.
Pubblicazione: (2023)
Selective, Interpretable, and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
di: Ilic, Filip, et al.
Pubblicazione: (2024)
di: Ilic, Filip, et al.
Pubblicazione: (2024)
ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography
di: Ghamary, Yeganeh, et al.
Pubblicazione: (2025)
di: Ghamary, Yeganeh, et al.
Pubblicazione: (2025)
Bridging Domains with Approximately Shared Features
di: Zhong, Ziliang Samuel, et al.
Pubblicazione: (2024)
di: Zhong, Ziliang Samuel, et al.
Pubblicazione: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
di: Guimard, Quentin, et al.
Pubblicazione: (2026)
di: Guimard, Quentin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024) -
Towards Unifying Interpretability and Control: Evaluation via Intervention
di: Bhalla, Usha, et al.
Pubblicazione: (2024) -
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025) -
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025) -
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
di: Srinivas, Suraj, et al.
Pubblicazione: (2023)