Can LLMs Explain Themselves Counterfactually?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dehghanighobadi, Zahra, Fischer, Asja, Zafar, Muhammad Bilal |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
par: Dehghanighobadi, Zahra, et autres
Publié: (2026)
par: Dehghanighobadi, Zahra, et autres
Publié: (2026)
LLMs Can Teach Themselves to Better Predict the Future
par: Turtel, Benjamin, et autres
Publié: (2025)
par: Turtel, Benjamin, et autres
Publié: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
par: Tie, Guiyao, et autres
Publié: (2025)
par: Tie, Guiyao, et autres
Publié: (2025)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
par: Wang, Qianli, et autres
Publié: (2026)
par: Wang, Qianli, et autres
Publié: (2026)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
par: Tan, Nelvin, et autres
Publié: (2025)
par: Tan, Nelvin, et autres
Publié: (2025)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
par: Neumann, Anna, et autres
Publié: (2025)
par: Neumann, Anna, et autres
Publié: (2025)
The Impact of Inference Acceleration on Bias of LLMs
par: Kirsten, Elisabeth, et autres
Publié: (2024)
par: Kirsten, Elisabeth, et autres
Publié: (2024)
On Early Detection of Hallucinations in Factual Question Answering
par: Snyder, Ben, et autres
Publié: (2023)
par: Snyder, Ben, et autres
Publié: (2023)
Looking Inward: Language Models Can Learn About Themselves by Introspection
par: Binder, Felix J, et autres
Publié: (2024)
par: Binder, Felix J, et autres
Publié: (2024)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
par: Xu, Wenda, et autres
Publié: (2025)
par: Xu, Wenda, et autres
Publié: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
par: Yang, Gao, et autres
Publié: (2025)
par: Yang, Gao, et autres
Publié: (2025)
ItD: Large Language Models Can Teach Themselves Induction through Deduction
par: Sun, Wangtao, et autres
Publié: (2024)
par: Sun, Wangtao, et autres
Publié: (2024)
Kantian-Utilitarian XAI: Meta-Explained
par: Atf, Zahra, et autres
Publié: (2025)
par: Atf, Zahra, et autres
Publié: (2025)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
par: Gogoulou, Evangelia, et autres
Publié: (2025)
par: Gogoulou, Evangelia, et autres
Publié: (2025)
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
par: Bilal, Ahsan, et autres
Publié: (2025)
par: Bilal, Ahsan, et autres
Publié: (2025)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
par: Zelikman, Eric, et autres
Publié: (2024)
par: Zelikman, Eric, et autres
Publié: (2024)
Do LLM hallucination detectors suffer from low-resource effect?
par: Datta, Debtanu, et autres
Publié: (2026)
par: Datta, Debtanu, et autres
Publié: (2026)
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
par: Kim, Kyung-Hoon
Publié: (2025)
par: Kim, Kyung-Hoon
Publié: (2025)
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
par: Ji, Jiazhou, et autres
Publié: (2025)
par: Ji, Jiazhou, et autres
Publié: (2025)
Language Models can Evaluate Themselves via Probability Discrepancy
par: Xia, Tingyu, et autres
Publié: (2024)
par: Xia, Tingyu, et autres
Publié: (2024)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
par: Nguyen, Van Bach, et autres
Publié: (2024)
par: Nguyen, Van Bach, et autres
Publié: (2024)
Using LLMs to identify features of personal and professional skills in an open-response situational judgment test
par: Walsh, Cole, et autres
Publié: (2025)
par: Walsh, Cole, et autres
Publié: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
par: Ghosh, Rajarshi, et autres
Publié: (2025)
par: Ghosh, Rajarshi, et autres
Publié: (2025)
Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
par: Alkaeed, Mahdi, et autres
Publié: (2026)
par: Alkaeed, Mahdi, et autres
Publié: (2026)
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
par: Muhaimin, Sheikh Samit, et autres
Publié: (2025)
par: Muhaimin, Sheikh Samit, et autres
Publié: (2025)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
par: Puerto, Haritz, et autres
Publié: (2026)
par: Puerto, Haritz, et autres
Publié: (2026)
LLMs for Explainable AI: A Comprehensive Survey
par: Bilal, Ahsan, et autres
Publié: (2025)
par: Bilal, Ahsan, et autres
Publié: (2025)
Reference-based Metrics Disprove Themselves in Question Generation
par: Nguyen, Bang, et autres
Publié: (2024)
par: Nguyen, Bang, et autres
Publié: (2024)
Natural Language Processing for Analyzing Electronic Health Records and Clinical Notes in Cancer Research: A Review
par: Bilal, Muhammad, et autres
Publié: (2024)
par: Bilal, Muhammad, et autres
Publié: (2024)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
par: Toker, Gilat, et autres
Publié: (2026)
par: Toker, Gilat, et autres
Publié: (2026)
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
par: Chen, Wenting, et autres
Publié: (2026)
par: Chen, Wenting, et autres
Publié: (2026)
Can LLMs Ask Good Questions?
par: Zhang, Yueheng, et autres
Publié: (2025)
par: Zhang, Yueheng, et autres
Publié: (2025)
Can GNN be Good Adapter for LLMs?
par: Huang, Xuanwen, et autres
Publié: (2024)
par: Huang, Xuanwen, et autres
Publié: (2024)
Can LLMs Capture Human Preferences?
par: Goli, Ali, et autres
Publié: (2023)
par: Goli, Ali, et autres
Publié: (2023)
Are Today's LLMs Ready to Explain Well-Being Concepts?
par: Jiang, Bohan, et autres
Publié: (2025)
par: Jiang, Bohan, et autres
Publié: (2025)
Can LLMs perform structured graph reasoning?
par: Agrawal, Palaash, et autres
Publié: (2024)
par: Agrawal, Palaash, et autres
Publié: (2024)
Can We Locate and Prevent Stereotypes in LLMs?
par: D'Souza, Alex
Publié: (2026)
par: D'Souza, Alex
Publié: (2026)
Can LLMs Perceive Time? An Empirical Investigation
par: Garikaparthi, Aniketh
Publié: (2026)
par: Garikaparthi, Aniketh
Publié: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
par: Hamman, Faisal, et autres
Publié: (2025)
par: Hamman, Faisal, et autres
Publié: (2025)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
par: Yang, Zhuonan, et autres
Publié: (2026)
par: Yang, Zhuonan, et autres
Publié: (2026)
Documents similaires
-
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
par: Dehghanighobadi, Zahra, et autres
Publié: (2026) -
LLMs Can Teach Themselves to Better Predict the Future
par: Turtel, Benjamin, et autres
Publié: (2025) -
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
par: Tie, Guiyao, et autres
Publié: (2025) -
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
par: Wang, Qianli, et autres
Publié: (2026) -
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
par: Tan, Nelvin, et autres
Publié: (2025)