Saved in:
| Main Authors: | Krishna, Satyapriya, Han, Tessa, Gu, Alex, Wu, Steven, Jabbari, Shahin, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2202.01602 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
On the Trade-offs between Adversarial Robustness and Actionable Explanations
by: Krishna, Satyapriya, et al.
Published: (2023)
by: Krishna, Satyapriya, et al.
Published: (2023)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
Characterizing Data Point Vulnerability via Average-Case Robustness
by: Han, Tessa, et al.
Published: (2023)
by: Han, Tessa, et al.
Published: (2023)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024)
by: Ley, Dan, et al.
Published: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)
by: Oesterling, Alex, et al.
Published: (2023)
A Study on the Calibration of In-context Learning
by: Zhang, Hanlin, et al.
Published: (2023)
by: Zhang, Hanlin, et al.
Published: (2023)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
by: Wu, Tianyu, et al.
Published: (2026)
by: Wu, Tianyu, et al.
Published: (2026)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
Matching Problems to Solutions: An Explainable Way of Solving Machine Learning Problems
by: Saleh, Lokman, et al.
Published: (2024)
by: Saleh, Lokman, et al.
Published: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
A Comprehensive Perspective on Explainable AI across the Machine Learning Workflow
by: Paterakis, George, et al.
Published: (2025)
by: Paterakis, George, et al.
Published: (2025)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023)
by: Bhalla, Usha, et al.
Published: (2023)
Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers
by: Du, Yuxuan, et al.
Published: (2025)
by: Du, Yuxuan, et al.
Published: (2025)
Towards Interpretable Soft Prompts
by: Patel, Oam, et al.
Published: (2025)
by: Patel, Oam, et al.
Published: (2025)
Explainable Machine Learning for ICU Readmission Prediction
by: de Sá, Alex G. C., et al.
Published: (2023)
by: de Sá, Alex G. C., et al.
Published: (2023)
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024)
by: Kumar, Aounon, et al.
Published: (2024)
Mitigating Spurious Correlations via Disagreement Probability
by: Han, Hyeonggeun, et al.
Published: (2024)
by: Han, Hyeonggeun, et al.
Published: (2024)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Verifying Machine Unlearning with Explainable AI
by: Vidal, Àlex Pujol, et al.
Published: (2024)
by: Vidal, Àlex Pujol, et al.
Published: (2024)
DIVE: Subgraph Disagreement for Graph Out-of-Distribution Generalization
by: Sun, Xin, et al.
Published: (2024)
by: Sun, Xin, et al.
Published: (2024)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving
by: Villuri, Gnaneswar, et al.
Published: (2024)
by: Villuri, Gnaneswar, et al.
Published: (2024)
On the Relationship Between Interpretability and Explainability in Machine Learning
by: Leblanc, Benjamin, et al.
Published: (2023)
by: Leblanc, Benjamin, et al.
Published: (2023)
Investigating the Duality of Interpretability and Explainability in Machine Learning
by: Garouani, Moncef, et al.
Published: (2025)
by: Garouani, Moncef, et al.
Published: (2025)
OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning
by: Ma, Jiaqi, et al.
Published: (2024)
by: Ma, Jiaqi, et al.
Published: (2024)
Similar Items
-
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025) -
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023) -
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024) -
On the Trade-offs between Adversarial Robustness and Actionable Explanations
by: Krishna, Satyapriya, et al.
Published: (2023) -
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)