Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Unlu, Eren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
von: Unlu, Eren
Veröffentlicht: (2026)
von: Unlu, Eren
Veröffentlicht: (2026)
Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts
von: Unlu, Eren
Veröffentlicht: (2026)
von: Unlu, Eren
Veröffentlicht: (2026)
Geotokens and Geotransformers
von: Unlu, Eren
Veröffentlicht: (2024)
von: Unlu, Eren
Veröffentlicht: (2024)
Architecting Trust in Artificial Epistemic Agents
von: Marchal, Nahema, et al.
Veröffentlicht: (2026)
von: Marchal, Nahema, et al.
Veröffentlicht: (2026)
Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
von: Manchingal, Shireen Kudukkil, et al.
Veröffentlicht: (2025)
von: Manchingal, Shireen Kudukkil, et al.
Veröffentlicht: (2025)
Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know
von: Manchingal, Shireen Kudukkil
Veröffentlicht: (2025)
von: Manchingal, Shireen Kudukkil
Veröffentlicht: (2025)
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
von: Xu, Binyan, et al.
Veröffentlicht: (2026)
von: Xu, Binyan, et al.
Veröffentlicht: (2026)
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
von: Li, Xiaoxiao
Veröffentlicht: (2026)
von: Li, Xiaoxiao
Veröffentlicht: (2026)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees
von: Gui, Yu, et al.
Veröffentlicht: (2024)
von: Gui, Yu, et al.
Veröffentlicht: (2024)
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
von: Stoisser, Josefa Lia, et al.
Veröffentlicht: (2025)
von: Stoisser, Josefa Lia, et al.
Veröffentlicht: (2025)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
von: Shao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Shao, Jiaqi, et al.
Veröffentlicht: (2025)
Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025)
von: Liu, Grace, et al.
Veröffentlicht: (2025)
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
von: Wang, Qingni, et al.
Veröffentlicht: (2026)
von: Wang, Qingni, et al.
Veröffentlicht: (2026)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
von: Hao, Chenjie, et al.
Veröffentlicht: (2026)
von: Hao, Chenjie, et al.
Veröffentlicht: (2026)
Trust & Safety of LLMs and LLMs in Trust & Safety
von: You, Doohee, et al.
Veröffentlicht: (2024)
von: You, Doohee, et al.
Veröffentlicht: (2024)
Do LLMs Know When to Flip a Coin? Strategic Randomization through Reasoning and Experience
von: Yang, Lingyu
Veröffentlicht: (2025)
von: Yang, Lingyu
Veröffentlicht: (2025)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
Knowing When to Stop: Delay-Adaptive Spiking Neural Network Classifiers with Reliability Guarantees
von: Chen, Jiechen, et al.
Veröffentlicht: (2023)
von: Chen, Jiechen, et al.
Veröffentlicht: (2023)
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
von: Trinh, Tu, et al.
Veröffentlicht: (2026)
von: Trinh, Tu, et al.
Veröffentlicht: (2026)
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
von: Maynard, Andrew D.
Veröffentlicht: (2026)
von: Maynard, Andrew D.
Veröffentlicht: (2026)
Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets
von: Nel, Lukas
Veröffentlicht: (2025)
von: Nel, Lukas
Veröffentlicht: (2025)
More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries
von: Song, Hongwen, et al.
Veröffentlicht: (2026)
von: Song, Hongwen, et al.
Veröffentlicht: (2026)
On the Performance of LLMs for Real Estate Appraisal
von: Geerts, Margot, et al.
Veröffentlicht: (2025)
von: Geerts, Margot, et al.
Veröffentlicht: (2025)
When Agents Persuade: Rhetoric Generation and Mitigation in LLMs
von: Jose, Julia, et al.
Veröffentlicht: (2026)
von: Jose, Julia, et al.
Veröffentlicht: (2026)
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2026)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
von: Zhou, Youchao, et al.
Veröffentlicht: (2025)
von: Zhou, Youchao, et al.
Veröffentlicht: (2025)
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
von: Sun, Zhaoyue, et al.
Veröffentlicht: (2026)
von: Sun, Zhaoyue, et al.
Veröffentlicht: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
The Confidence Paradox: Can LLM Know When It's Wrong
von: Tripathi, Sahil, et al.
Veröffentlicht: (2025)
von: Tripathi, Sahil, et al.
Veröffentlicht: (2025)
Agents Need Not Know Their Purpose
von: Garcia, Paulo
Veröffentlicht: (2024)
von: Garcia, Paulo
Veröffentlicht: (2024)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
von: Wang, Su, et al.
Veröffentlicht: (2026)
von: Wang, Su, et al.
Veröffentlicht: (2026)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates
von: Zou, Rui, et al.
Veröffentlicht: (2024)
von: Zou, Rui, et al.
Veröffentlicht: (2024)
KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
von: Pan, Jinhao, et al.
Veröffentlicht: (2026)
von: Pan, Jinhao, et al.
Veröffentlicht: (2026)
Epistemic Skills: Reasoning about Knowledge and Oblivion
von: Liang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Liang, Xiaolong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
von: Unlu, Eren
Veröffentlicht: (2026) -
Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts
von: Unlu, Eren
Veröffentlicht: (2026) -
Geotokens and Geotransformers
von: Unlu, Eren
Veröffentlicht: (2024) -
Architecting Trust in Artificial Epistemic Agents
von: Marchal, Nahema, et al.
Veröffentlicht: (2026) -
Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
von: Manchingal, Shireen Kudukkil, et al.
Veröffentlicht: (2025)