On the Trade-offs between Adversarial Robustness and Actionable Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Krishna, Satyapriya, Agarwal, Chirag, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
by: Krishna, Satyapriya, et al.
Published: (2022)
by: Krishna, Satyapriya, et al.
Published: (2022)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Characterizing Data Point Vulnerability via Average-Case Robustness
by: Han, Tessa, et al.
Published: (2023)
by: Han, Tessa, et al.
Published: (2023)
Explaining the Model, Protecting Your Data: Revealing and Mitigating the Data Privacy Risks of Post-Hoc Model Explanations via Membership Inference
by: Huang, Catherine, et al.
Published: (2024)
by: Huang, Catherine, et al.
Published: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
by: Agarwal, Chirag, et al.
Published: (2024)
by: Agarwal, Chirag, et al.
Published: (2024)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
Transparent Trade-offs between Properties of Explanations
by: Tadesse, Hiwot Belay, et al.
Published: (2024)
by: Tadesse, Hiwot Belay, et al.
Published: (2024)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023)
by: Bhalla, Usha, et al.
Published: (2023)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Quantifying Explanation Quality in Graph Neural Networks using Out-of-Distribution Generalization
by: Zhang, Ding, et al.
Published: (2026)
by: Zhang, Ding, et al.
Published: (2026)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
by: Liu, Yanchen, et al.
Published: (2023)
by: Liu, Yanchen, et al.
Published: (2023)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
by: Badrinath, Charumathi, et al.
Published: (2024)
by: Badrinath, Charumathi, et al.
Published: (2024)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
by: Srinivas, Suraj, et al.
Published: (2023)
by: Srinivas, Suraj, et al.
Published: (2023)
Generalized Group Data Attribution
by: Ley, Dan, et al.
Published: (2024)
by: Ley, Dan, et al.
Published: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Rethinking Explainability in the Era of Multimodal AI
by: Agarwal, Chirag
Published: (2025)
by: Agarwal, Chirag
Published: (2025)
Learning Actionable Counterfactual Explanations in Large State Spaces
by: Naggita, Keziah, et al.
Published: (2024)
by: Naggita, Keziah, et al.
Published: (2024)
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
by: Tanneru, Sree Harsha, et al.
Published: (2024)
by: Tanneru, Sree Harsha, et al.
Published: (2024)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Causal Explanation of Concept Drift -- A Truly Actionable Approach
by: Komnick, David, et al.
Published: (2025)
by: Komnick, David, et al.
Published: (2025)
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025)
by: Khalaf, Hadi, et al.
Published: (2025)
Rethinking Invariance Regularization in Adversarial Training to Improve Robustness-Accuracy Trade-off
by: Waseda, Futa, et al.
Published: (2024)
by: Waseda, Futa, et al.
Published: (2024)
Counterfactual Training: Teaching Models Plausible and Actionable Explanations
by: Altmeyer, Patrick, et al.
Published: (2026)
by: Altmeyer, Patrick, et al.
Published: (2026)
PUPAE: Intuitive and Actionable Explanations for Time Series Anomalies
by: Der, Audrey, et al.
Published: (2024)
by: Der, Audrey, et al.
Published: (2024)
Improving the Evaluation and Actionability of Explanation Methods for Multivariate Time Series Classification
by: Serramazza, Davide Italo, et al.
Published: (2024)
by: Serramazza, Davide Italo, et al.
Published: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Similar Items
-
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023) -
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024) -
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022) -
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
by: Krishna, Satyapriya, et al.
Published: (2022) -
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)