Prompt-Counterfactual Explanations for Generative AI System Behavior
Fuente:
arXiv
Salvato in:
| Autori principali: | Goethals, Sofie, Provost, Foster, Sedoc, João |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
Counterfactual Explanations for Hypergraph Neural Networks
di: Veglianti, Fabiano, et al.
Pubblicazione: (2026)
di: Veglianti, Fabiano, et al.
Pubblicazione: (2026)
Procedural Fairness via Group Counterfactual Explanation
di: Popoola, Gideon, et al.
Pubblicazione: (2026)
di: Popoola, Gideon, et al.
Pubblicazione: (2026)
Properties and Challenges of LLM-Generated Explanations
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
di: Wilming, Rick, et al.
Pubblicazione: (2024)
di: Wilming, Rick, et al.
Pubblicazione: (2024)
Large Human Language Models: A Need and the Challenges
di: Soni, Nikita, et al.
Pubblicazione: (2023)
di: Soni, Nikita, et al.
Pubblicazione: (2023)
Understanding and Mitigating Risks of Generative AI in Financial Services
di: Gehrmann, Sebastian, et al.
Pubblicazione: (2025)
di: Gehrmann, Sebastian, et al.
Pubblicazione: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
di: Ren, Richard, et al.
Pubblicazione: (2025)
di: Ren, Richard, et al.
Pubblicazione: (2025)
M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
di: Jeong, Jihoon
Pubblicazione: (2026)
di: Jeong, Jihoon
Pubblicazione: (2026)
KTCF: Actionable Recourse in Knowledge Tracing via Counterfactual Explanations for Education
di: Kim, Woojin, et al.
Pubblicazione: (2026)
di: Kim, Woojin, et al.
Pubblicazione: (2026)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
di: Shin, Kwan Soo
Pubblicazione: (2026)
di: Shin, Kwan Soo
Pubblicazione: (2026)
TABCF: Counterfactual Explanations for Tabular Data Using a Transformer-Based VAE
di: Panagiotou, Emmanouil, et al.
Pubblicazione: (2024)
di: Panagiotou, Emmanouil, et al.
Pubblicazione: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
EigenBench: A Comparative Behavioral Measure of Value Alignment
di: Chang, Jonathn, et al.
Pubblicazione: (2025)
di: Chang, Jonathn, et al.
Pubblicazione: (2025)
On the Definition and Detection of Cherry-Picking in Counterfactual Explanations
di: Hinns, James, et al.
Pubblicazione: (2026)
di: Hinns, James, et al.
Pubblicazione: (2026)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
di: Han, Pengrui, et al.
Pubblicazione: (2025)
di: Han, Pengrui, et al.
Pubblicazione: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
di: Chen, Ruxiao, et al.
Pubblicazione: (2025)
di: Chen, Ruxiao, et al.
Pubblicazione: (2025)
AI-AI Bias: large language models favor communications generated by large language models
di: Laurito, Walter, et al.
Pubblicazione: (2024)
di: Laurito, Walter, et al.
Pubblicazione: (2024)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
di: Hamman, Faisal, et al.
Pubblicazione: (2023)
di: Hamman, Faisal, et al.
Pubblicazione: (2023)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
di: Schroeder, Daniel Thilo, et al.
Pubblicazione: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
Managing extreme AI risks amid rapid progress
di: Bengio, Yoshua, et al.
Pubblicazione: (2023)
di: Bengio, Yoshua, et al.
Pubblicazione: (2023)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
di: Tang, Xiangru, et al.
Pubblicazione: (2024)
di: Tang, Xiangru, et al.
Pubblicazione: (2024)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
di: Bai, Xiaoyan, et al.
Pubblicazione: (2025)
di: Bai, Xiaoyan, et al.
Pubblicazione: (2025)
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
di: Andric, Sandro
Pubblicazione: (2025)
di: Andric, Sandro
Pubblicazione: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
Impacts of Racial Bias in Historical Training Data for News AI
di: Bhargava, Rahul, et al.
Pubblicazione: (2025)
di: Bhargava, Rahul, et al.
Pubblicazione: (2025)
Generative AI Security: Challenges and Countermeasures
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
di: Schoenegger, Philipp, et al.
Pubblicazione: (2024)
di: Schoenegger, Philipp, et al.
Pubblicazione: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
di: Yang, Chao, et al.
Pubblicazione: (2024)
di: Yang, Chao, et al.
Pubblicazione: (2024)
Frontier AI systems have surpassed the self-replicating red line
di: Pan, Xudong, et al.
Pubblicazione: (2024)
di: Pan, Xudong, et al.
Pubblicazione: (2024)
On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
di: Geng, Mingmeng, et al.
Pubblicazione: (2025)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2023)
di: Zhang, Yue, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
di: Hamman, Faisal, et al.
Pubblicazione: (2025) -
Counterfactual Explanations for Hypergraph Neural Networks
di: Veglianti, Fabiano, et al.
Pubblicazione: (2026) -
Procedural Fairness via Group Counterfactual Explanation
di: Popoola, Gideon, et al.
Pubblicazione: (2026) -
Properties and Challenges of LLM-Generated Explanations
di: Kunz, Jenny, et al.
Pubblicazione: (2024) -
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
di: Wilming, Rick, et al.
Pubblicazione: (2024)