Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
Fuente:
arXiv
Guardado en:
| Autor principal: | Wang, Yufeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
por: Zhang, Mingxuan, et al.
Publicado: (2026)
por: Zhang, Mingxuan, et al.
Publicado: (2026)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
por: Sirdeshmukh, Ved, et al.
Publicado: (2026)
por: Sirdeshmukh, Ved, et al.
Publicado: (2026)
Reasoning Models Don't Always Say What They Think
por: Chen, Yanda, et al.
Publicado: (2025)
por: Chen, Yanda, et al.
Publicado: (2025)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
por: Puerto, Haritz, et al.
Publicado: (2026)
por: Puerto, Haritz, et al.
Publicado: (2026)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
por: Wu, Yilin, et al.
Publicado: (2025)
por: Wu, Yilin, et al.
Publicado: (2025)
What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
por: Szeider, Stefan
Publicado: (2025)
por: Szeider, Stefan
Publicado: (2025)
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
por: Cheng, Liang, et al.
Publicado: (2026)
por: Cheng, Liang, et al.
Publicado: (2026)
Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
por: Salgado, Henry, et al.
Publicado: (2025)
por: Salgado, Henry, et al.
Publicado: (2025)
AlgBench: To What Extent Do Large Reasoning Models Understand Algorithms?
por: Sun, Henan, et al.
Publicado: (2026)
por: Sun, Henan, et al.
Publicado: (2026)
Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?
por: Liu, Xiaoze, et al.
Publicado: (2026)
por: Liu, Xiaoze, et al.
Publicado: (2026)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
por: Alon, Bar, et al.
Publicado: (2026)
por: Alon, Bar, et al.
Publicado: (2026)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
por: Xu, Ruijie, et al.
Publicado: (2024)
por: Xu, Ruijie, et al.
Publicado: (2024)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
por: Yuan, Wenhao, et al.
Publicado: (2026)
por: Yuan, Wenhao, et al.
Publicado: (2026)
What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding
por: Liu, Siyuan, et al.
Publicado: (2026)
por: Liu, Siyuan, et al.
Publicado: (2026)
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
por: Mannekote, Amogh, et al.
Publicado: (2025)
por: Mannekote, Amogh, et al.
Publicado: (2025)
What Deserves Memory: Adaptive Memory Distillation for LLM Agents
por: Ma, Wenquan, et al.
Publicado: (2025)
por: Ma, Wenquan, et al.
Publicado: (2025)
What Is AI Safety? What Do We Want It to Be?
por: Harding, Jacqueline, et al.
Publicado: (2025)
por: Harding, Jacqueline, et al.
Publicado: (2025)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
por: Shao, Jiaqi, et al.
Publicado: (2025)
por: Shao, Jiaqi, et al.
Publicado: (2025)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
por: Hoang, Huy, et al.
Publicado: (2025)
por: Hoang, Huy, et al.
Publicado: (2025)
What Would an LLM Do? Evaluating Large Language Models for Policymaking to Alleviate Homelessness
por: Coz, Pierre Le, et al.
Publicado: (2025)
por: Coz, Pierre Le, et al.
Publicado: (2025)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
por: Shen, Xu, et al.
Publicado: (2025)
por: Shen, Xu, et al.
Publicado: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
por: Mittal, Avni, et al.
Publicado: (2026)
por: Mittal, Avni, et al.
Publicado: (2026)
When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
por: Yamin, Khurram, et al.
Publicado: (2026)
por: Yamin, Khurram, et al.
Publicado: (2026)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
por: Yan, Lewen, et al.
Publicado: (2025)
por: Yan, Lewen, et al.
Publicado: (2025)
AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning
por: Fu, Suzhong, et al.
Publicado: (2026)
por: Fu, Suzhong, et al.
Publicado: (2026)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
por: Zhao, Tianzhe, et al.
Publicado: (2026)
por: Zhao, Tianzhe, et al.
Publicado: (2026)
What Do Learned Models Measure?
por: Žliobaitė, Indrė
Publicado: (2026)
por: Žliobaitė, Indrė
Publicado: (2026)
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents
por: Xu, Rui, et al.
Publicado: (2025)
por: Xu, Rui, et al.
Publicado: (2025)
Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents
por: Khanzadeh, Sourena
Publicado: (2026)
por: Khanzadeh, Sourena
Publicado: (2026)
Closing Reasoning Gaps in Clinical Agents with Differential Reasoning Learning
por: Liu, Jinsong, et al.
Publicado: (2026)
por: Liu, Jinsong, et al.
Publicado: (2026)
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
por: Jung, Chaeyoung, et al.
Publicado: (2026)
por: Jung, Chaeyoung, et al.
Publicado: (2026)
Time-Scaling Is What Agents Need Now
por: Liu, Zhi, et al.
Publicado: (2026)
por: Liu, Zhi, et al.
Publicado: (2026)
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?
por: Yin, Xiaoyun, et al.
Publicado: (2025)
por: Yin, Xiaoyun, et al.
Publicado: (2025)
What Do Latent Action Models Actually Learn?
por: Zhang, Chuheng, et al.
Publicado: (2025)
por: Zhang, Chuheng, et al.
Publicado: (2025)
What Do AI-Generated Images Want?
por: Wasielewski, Amanda
Publicado: (2025)
por: Wasielewski, Amanda
Publicado: (2025)
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
por: Li, Yubo, et al.
Publicado: (2026)
por: Li, Yubo, et al.
Publicado: (2026)
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
por: Camassa, Carolina, et al.
Publicado: (2026)
por: Camassa, Carolina, et al.
Publicado: (2026)
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
por: Lin, Dongding, et al.
Publicado: (2026)
por: Lin, Dongding, et al.
Publicado: (2026)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
por: Parikh, Aditya, et al.
Publicado: (2026)
por: Parikh, Aditya, et al.
Publicado: (2026)
Mamba-SSM with LLM Reasoning for Feature Selection: Faithfulness-Aware Biomarker Discovery
por: Balan, Pushpa Kumar, et al.
Publicado: (2026)
por: Balan, Pushpa Kumar, et al.
Publicado: (2026)
Ejemplares similares
-
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
por: Zhang, Mingxuan, et al.
Publicado: (2026) -
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
por: Sirdeshmukh, Ved, et al.
Publicado: (2026) -
Reasoning Models Don't Always Say What They Think
por: Chen, Yanda, et al.
Publicado: (2025) -
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
por: Puerto, Haritz, et al.
Publicado: (2026) -
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
por: Wu, Yilin, et al.
Publicado: (2025)