LLM-Generated Black-box Explanations Can Be Adversarially Helpful
Fuente:
arXiv
Saved in:
| Main Authors: | Ajwani, Rohan, Javaji, Shashidhar Reddy, Rudzicz, Frank, Zhu, Zining |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
by: Javaji, Shashidhar Reddy, et al.
Published: (2024)
by: Javaji, Shashidhar Reddy, et al.
Published: (2024)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)
by: Qiu, Pengshuo, et al.
Published: (2024)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
Understanding Language Model Circuits through Knowledge Editing
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
VERBA: Verbalizing Model Differences Using Large Language Models
by: Doda, Shravan, et al.
Published: (2025)
by: Doda, Shravan, et al.
Published: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
by: Jaipersaud, Brandon, et al.
Published: (2024)
by: Jaipersaud, Brandon, et al.
Published: (2024)
Situated Natural Language Explanations
by: Zhu, Zining, et al.
Published: (2023)
by: Zhu, Zining, et al.
Published: (2023)
ACCORD: Closing the Commonsense Measurability Gap
by: Roewer-Després, François, et al.
Published: (2024)
by: Roewer-Després, François, et al.
Published: (2024)
Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation
by: Cui, Huizi, et al.
Published: (2026)
by: Cui, Huizi, et al.
Published: (2026)
LLM Library Learning Fails: A LEGO-Prover Case Study
by: Berlot-Attwell, Ian, et al.
Published: (2025)
by: Berlot-Attwell, Ian, et al.
Published: (2025)
Automated ensemble method for pediatric brain tumor segmentation
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
by: Wang, Haojin, et al.
Published: (2025)
by: Wang, Haojin, et al.
Published: (2025)
Length Controlled Generation for Black-box LLMs
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
by: Huang, Xi Yu, et al.
Published: (2024)
by: Huang, Xi Yu, et al.
Published: (2024)
Auxiliary Knowledge-Induced Learning for Automatic Multi-Label Medical Document Classification
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
Exploring the features used for summary evaluation by Human and GPT
by: Sadeghi, Zahra, et al.
Published: (2025)
by: Sadeghi, Zahra, et al.
Published: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024)
by: Berlot-Attwell, Ian, et al.
Published: (2024)
Multi-stage Retrieve and Re-rank Model for Automatic Medical Coding Recommendation
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
How Much Can RAG Help the Reasoning of LLM?
by: Liu, Jingyu, et al.
Published: (2024)
by: Liu, Jingyu, et al.
Published: (2024)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
by: Fragkathoulas, Christos, et al.
Published: (2024)
by: Fragkathoulas, Christos, et al.
Published: (2024)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
by: Zhou, Xiaotian, et al.
Published: (2026)
by: Zhou, Xiaotian, et al.
Published: (2026)
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Graph-tree Fusion Model with Bidirectional Information Propagation for Long Document Classification
by: Roy, Sudipta Singha, et al.
Published: (2024)
by: Roy, Sudipta Singha, et al.
Published: (2024)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Question Generation for Assessing Early Literacy Reading Comprehension
by: Yang, Xiaocheng, et al.
Published: (2025)
by: Yang, Xiaocheng, et al.
Published: (2025)
X-MuTeST: A Multilingual Benchmark for Explainable Hate Speech Detection and A Novel LLM-consulted Explanation Framework
by: Rehman, Mohammad Zia Ur, et al.
Published: (2026)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2026)
Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
by: Zhao, Yunxiao, et al.
Published: (2025)
by: Zhao, Yunxiao, et al.
Published: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
Adversarial Attack for Explanation Robustness of Rationalization Models
by: Zhang, Yuankai, et al.
Published: (2024)
by: Zhang, Yuankai, et al.
Published: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
by: Lin, Zhen, et al.
Published: (2023)
by: Lin, Zhen, et al.
Published: (2023)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
by: Limpijankit, Marvin, et al.
Published: (2025)
by: Limpijankit, Marvin, et al.
Published: (2025)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
Similar Items
-
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
by: Javaji, Shashidhar Reddy, et al.
Published: (2024) -
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
by: Ajwani, Rohan Deepak, et al.
Published: (2024) -
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025) -
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
by: Javaji, Shashidhar Reddy, et al.
Published: (2025) -
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)