LLM Self-Explanations Fail Semantic Invariance
Fuente:
arXiv
Salvato in:
| Autore principale: | Szeider, Stefan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ASP-Bench: From Natural Language to Logic Programs
di: Szeider, Stefan
Pubblicazione: (2026)
di: Szeider, Stefan
Pubblicazione: (2026)
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
di: Szeider, Stefan
Pubblicazione: (2026)
di: Szeider, Stefan
Pubblicazione: (2026)
CP-Agent: Agentic Constraint Programming
di: Szeider, Stefan
Pubblicazione: (2025)
di: Szeider, Stefan
Pubblicazione: (2025)
MCP-Solver: Integrating Language Models with Constraint Programming Systems
di: Szeider, Stefan
Pubblicazione: (2024)
di: Szeider, Stefan
Pubblicazione: (2024)
Semantic Invariance in Agentic AI
di: de Zarzà, I., et al.
Pubblicazione: (2026)
di: de Zarzà, I., et al.
Pubblicazione: (2026)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
di: Cohen, Ruth, et al.
Pubblicazione: (2026)
di: Cohen, Ruth, et al.
Pubblicazione: (2026)
PBLean: Pseudo-Boolean Proof Certificates for Lean 4
di: Szeider, Stefan
Pubblicazione: (2026)
di: Szeider, Stefan
Pubblicazione: (2026)
What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
di: Szeider, Stefan
Pubblicazione: (2025)
di: Szeider, Stefan
Pubblicazione: (2025)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
Support-Contra Asymmetry in LLM Explanations
di: Patil, Avinash
Pubblicazione: (2025)
di: Patil, Avinash
Pubblicazione: (2025)
Anchored Alignment for Self-Explanations Enhancement
di: Villa-Arenas, Luis Felipe, et al.
Pubblicazione: (2024)
di: Villa-Arenas, Luis Felipe, et al.
Pubblicazione: (2024)
LLM-as-a-Judge for Time Series Explanations
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
STRUX: An LLM for Decision-Making with Structured Explanations
di: Lu, Yiming, et al.
Pubblicazione: (2024)
di: Lu, Yiming, et al.
Pubblicazione: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
di: Aghzal, Mohamed, et al.
Pubblicazione: (2026)
di: Aghzal, Mohamed, et al.
Pubblicazione: (2026)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
Self-Explanation in Social AI Agents
di: Basappa, Rhea, et al.
Pubblicazione: (2025)
di: Basappa, Rhea, et al.
Pubblicazione: (2025)
Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
Integrating Hierarchical Semantic into Iterative Generation Model for Entailment Tree Explanation
di: Wang, Qin, et al.
Pubblicazione: (2024)
di: Wang, Qin, et al.
Pubblicazione: (2024)
Learning Disentangled Semantic Spaces of Explanations via Invertible Neural Networks
di: Zhang, Yingji, et al.
Pubblicazione: (2023)
di: Zhang, Yingji, et al.
Pubblicazione: (2023)
Rethinking Invariance in In-context Learning
di: Fang, Lizhe, et al.
Pubblicazione: (2025)
di: Fang, Lizhe, et al.
Pubblicazione: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
di: Quan, Xin, et al.
Pubblicazione: (2025)
di: Quan, Xin, et al.
Pubblicazione: (2025)
Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
di: Jiang, Chunyang, et al.
Pubblicazione: (2025)
di: Jiang, Chunyang, et al.
Pubblicazione: (2025)
Distilling Text Style Transfer With Self-Explanation From LLMs
di: Zhang, Chiyu, et al.
Pubblicazione: (2024)
di: Zhang, Chiyu, et al.
Pubblicazione: (2024)
PLEX: Perturbation-free Local Explanations for LLM-Based Text Classification
di: Rahulamathavan, Yogachandran, et al.
Pubblicazione: (2025)
di: Rahulamathavan, Yogachandran, et al.
Pubblicazione: (2025)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
Simulated Ignorance Fails: A Systematic Study of LLM Behaviors on Forecasting Problems Before Model Knowledge Cutoff
di: Li, Zehan, et al.
Pubblicazione: (2026)
di: Li, Zehan, et al.
Pubblicazione: (2026)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
di: Cho, Yousang, et al.
Pubblicazione: (2025)
di: Cho, Yousang, et al.
Pubblicazione: (2025)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
di: Kim, Juyeon, et al.
Pubblicazione: (2024)
di: Kim, Juyeon, et al.
Pubblicazione: (2024)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
di: Chen, Zichen, et al.
Pubblicazione: (2023)
di: Chen, Zichen, et al.
Pubblicazione: (2023)
SNAPE-PM: Building and Utilizing Dynamic Partner Models for Adaptive Explanation Generation
di: Robrecht, Amelie S., et al.
Pubblicazione: (2025)
di: Robrecht, Amelie S., et al.
Pubblicazione: (2025)
How LLMs Fail to Support Fact-Checking
di: Proma, Adiba Mahbub, et al.
Pubblicazione: (2025)
di: Proma, Adiba Mahbub, et al.
Pubblicazione: (2025)
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
di: Yang, Jing, et al.
Pubblicazione: (2024)
di: Yang, Jing, et al.
Pubblicazione: (2024)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
di: Landesberg, Eddie
Pubblicazione: (2026)
di: Landesberg, Eddie
Pubblicazione: (2026)
The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective
di: Zahedzadeh, Ali, et al.
Pubblicazione: (2026)
di: Zahedzadeh, Ali, et al.
Pubblicazione: (2026)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2025)
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2025)
Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
di: Chen, Zichen, et al.
Pubblicazione: (2025)
di: Chen, Zichen, et al.
Pubblicazione: (2025)
LLM-Agnostic Semantic Representation Attack
di: Lian, Jiawei, et al.
Pubblicazione: (2026)
di: Lian, Jiawei, et al.
Pubblicazione: (2026)
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ASP-Bench: From Natural Language to Logic Programs
di: Szeider, Stefan
Pubblicazione: (2026) -
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
di: Szeider, Stefan
Pubblicazione: (2026) -
CP-Agent: Agentic Constraint Programming
di: Szeider, Stefan
Pubblicazione: (2025) -
MCP-Solver: Integrating Language Models with Constraint Programming Systems
di: Szeider, Stefan
Pubblicazione: (2024) -
Semantic Invariance in Agentic AI
di: de Zarzà, I., et al.
Pubblicazione: (2026)