Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shichang, Han, Tessa, Bhalla, Usha, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023)
by: Bhalla, Usha, et al.
Published: (2023)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024)
by: Ley, Dan, et al.
Published: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
by: Krishna, Satyapriya, et al.
Published: (2022)
by: Krishna, Satyapriya, et al.
Published: (2022)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
Characterizing Data Point Vulnerability via Average-Case Robustness
by: Han, Tessa, et al.
Published: (2023)
by: Han, Tessa, et al.
Published: (2023)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
by: Badrinath, Charumathi, et al.
Published: (2024)
by: Badrinath, Charumathi, et al.
Published: (2024)
Towards Interpretable Soft Prompts
by: Patel, Oam, et al.
Published: (2025)
by: Patel, Oam, et al.
Published: (2025)
Efficient Ensembles Improve Training Data Attribution
by: Deng, Junwei, et al.
Published: (2024)
by: Deng, Junwei, et al.
Published: (2024)
Correlation-Aware Feature Attribution Based Explainable AI
by: Sengupta, Poushali, et al.
Published: (2025)
by: Sengupta, Poushali, et al.
Published: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
DSAI: Unbiased and Interpretable Latent Feature Extraction for Data-Centric AI
by: Cho, Hyowon, et al.
Published: (2024)
by: Cho, Hyowon, et al.
Published: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Computational Copyright: Towards A Royalty Model for Music Generative AI
by: Deng, Junwei, et al.
Published: (2023)
by: Deng, Junwei, et al.
Published: (2023)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Towards User-Focused Research in Training Data Attribution for Human-Centered Explainable AI
by: Nguyen, Elisa, et al.
Published: (2024)
by: Nguyen, Elisa, et al.
Published: (2024)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
Towards Interpretable End-Stage Renal Disease (ESRD) Prediction: Utilizing Administrative Claims Data with Explainable AI Techniques
by: Li, Yubo, et al.
Published: (2024)
by: Li, Yubo, et al.
Published: (2024)
Transparent Adaptive Learning via Data-Centric Multimodal Explainable AI
by: Mosleh, Maryam, et al.
Published: (2025)
by: Mosleh, Maryam, et al.
Published: (2025)
Interpretable Representations in Explainable AI: From Theory to Practice
by: Sokol, Kacper, et al.
Published: (2020)
by: Sokol, Kacper, et al.
Published: (2020)
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
by: Chen, Jianhui, et al.
Published: (2026)
by: Chen, Jianhui, et al.
Published: (2026)
DataMaster: Data-Centric Autonomous AI Research
by: Du, Yaxin, et al.
Published: (2026)
by: Du, Yaxin, et al.
Published: (2026)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
Transparent AI: The Case for Interpretability and Explainability
by: Ramachandram, Dhanesh, et al.
Published: (2025)
by: Ramachandram, Dhanesh, et al.
Published: (2025)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
ABE: A Unified Framework for Robust and Faithful Attribution-Based Explainability
by: Zhu, Zhiyu, et al.
Published: (2025)
by: Zhu, Zhiyu, et al.
Published: (2025)
Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation
by: Wang, Dongjie, et al.
Published: (2025)
by: Wang, Dongjie, et al.
Published: (2025)
Unifying VXAI: A Systematic Review and Framework for the Evaluation of Explainable AI
by: Dembinsky, David, et al.
Published: (2025)
by: Dembinsky, David, et al.
Published: (2025)
A Survey on Data-Centric AI: Tabular Learning from Reinforcement Learning and Generative AI Perspective
by: Ying, Wangyang, et al.
Published: (2025)
by: Ying, Wangyang, et al.
Published: (2025)
Interpretable and Explainable Surrogate Modeling for Simulations: A State-of-the-Art Survey and Perspectives on Explainable AI for Decision-Making
by: Palar, Pramudita Satria, et al.
Published: (2026)
by: Palar, Pramudita Satria, et al.
Published: (2026)
xEEGNet: Towards Explainable AI in EEG Dementia Classification
by: Zanola, Andrea, et al.
Published: (2025)
by: Zanola, Andrea, et al.
Published: (2025)
Similar Items
-
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023) -
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024) -
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
by: Zhang, Shichang, et al.
Published: (2025) -
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024) -
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)