Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shichang, Du, Hongzhe, Ma, Jiaqi W., Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024)
by: Ley, Dan, et al.
Published: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
Efficient Ensembles Improve Training Data Attribution
by: Deng, Junwei, et al.
Published: (2024)
by: Deng, Junwei, et al.
Published: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Computational Copyright: Towards A Royalty Model for Music Generative AI
by: Deng, Junwei, et al.
Published: (2023)
by: Deng, Junwei, et al.
Published: (2023)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
by: Krishna, Satyapriya, et al.
Published: (2022)
by: Krishna, Satyapriya, et al.
Published: (2022)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)
by: Oesterling, Alex, et al.
Published: (2023)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
by: Wu, Tianyu, et al.
Published: (2026)
by: Wu, Tianyu, et al.
Published: (2026)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
by: Mundt, Martin, et al.
Published: (2025)
by: Mundt, Martin, et al.
Published: (2025)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
A Study on the Calibration of In-context Learning
by: Zhang, Hanlin, et al.
Published: (2023)
by: Zhang, Hanlin, et al.
Published: (2023)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023)
by: Bhalla, Usha, et al.
Published: (2023)
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
by: Hu, Pingbang, et al.
Published: (2025)
by: Hu, Pingbang, et al.
Published: (2025)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
by: Liu, Yanchen, et al.
Published: (2023)
by: Liu, Yanchen, et al.
Published: (2023)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
by: Yin, Junxi, et al.
Published: (2025)
by: Yin, Junxi, et al.
Published: (2025)
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024)
by: Kumar, Aounon, et al.
Published: (2024)
BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
by: Armstrong, Joss
Published: (2026)
by: Armstrong, Joss
Published: (2026)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
by: Buckley, Warren, et al.
Published: (2023)
by: Buckley, Warren, et al.
Published: (2023)
Comparing Credit Risk Estimates in the Gen-AI Era
by: Lavecchia, Nicola, et al.
Published: (2025)
by: Lavecchia, Nicola, et al.
Published: (2025)
Towards Interpretable Soft Prompts
by: Patel, Oam, et al.
Published: (2025)
by: Patel, Oam, et al.
Published: (2025)
Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
by: Dixit, Shrey, et al.
Published: (2025)
by: Dixit, Shrey, et al.
Published: (2025)
DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Correlation-Aware Feature Attribution Based Explainable AI
by: Sengupta, Poushali, et al.
Published: (2025)
by: Sengupta, Poushali, et al.
Published: (2025)
Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
by: Yang, Chih-Hsuan, et al.
Published: (2025)
by: Yang, Chih-Hsuan, et al.
Published: (2025)
Attribution Explanations for Deep Neural Networks: A Theoretical Perspective
by: Deng, Huiqi, et al.
Published: (2025)
by: Deng, Huiqi, et al.
Published: (2025)
On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
by: Widener, Michael, et al.
Published: (2025)
by: Widener, Michael, et al.
Published: (2025)
SETA: Statistical Fault Attribution for Compound AI Systems
by: Chowdhury, Sayak, et al.
Published: (2026)
by: Chowdhury, Sayak, et al.
Published: (2026)
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
by: Du, Yuexi, et al.
Published: (2026)
by: Du, Yuexi, et al.
Published: (2026)
Similar Items
-
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025) -
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024) -
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025) -
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024) -
Efficient Ensembles Improve Training Data Attribution
by: Deng, Junwei, et al.
Published: (2024)