Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Cheng, Liang, Cai, Mingsheng, Jiang, Jiuming, Mai, Luo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
por: Priyanshu, Aman, et al.
Publicado: (2026)
por: Priyanshu, Aman, et al.
Publicado: (2026)
What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding
por: Liu, Siyuan, et al.
Publicado: (2026)
por: Liu, Siyuan, et al.
Publicado: (2026)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
por: Shao, Jiaqi, et al.
Publicado: (2025)
por: Shao, Jiaqi, et al.
Publicado: (2025)
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
por: Cheng, Yize, et al.
Publicado: (2026)
por: Cheng, Yize, et al.
Publicado: (2026)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
por: P, Vedanta S, et al.
Publicado: (2026)
por: P, Vedanta S, et al.
Publicado: (2026)
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
por: Trinh, Tu, et al.
Publicado: (2026)
por: Trinh, Tu, et al.
Publicado: (2026)
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
por: Unlu, Eren
Publicado: (2026)
por: Unlu, Eren
Publicado: (2026)
Do Large Language Models Know What They Are Capable Of?
por: Barkan, Casey O., et al.
Publicado: (2025)
por: Barkan, Casey O., et al.
Publicado: (2025)
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
por: Du, Xiaodan, et al.
Publicado: (2023)
por: Du, Xiaodan, et al.
Publicado: (2023)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
por: Hoang, Huy, et al.
Publicado: (2025)
por: Hoang, Huy, et al.
Publicado: (2025)
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
por: Wang, Yufeng
Publicado: (2026)
por: Wang, Yufeng
Publicado: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
por: Pan, Wenbo, et al.
Publicado: (2025)
por: Pan, Wenbo, et al.
Publicado: (2025)
What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
por: Szeider, Stefan
Publicado: (2025)
por: Szeider, Stefan
Publicado: (2025)
From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets
por: Zhu, Taojie, et al.
Publicado: (2026)
por: Zhu, Taojie, et al.
Publicado: (2026)
SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning
por: Cai, Mingsheng, et al.
Publicado: (2025)
por: Cai, Mingsheng, et al.
Publicado: (2025)
Can AI Assistants Know What They Don't Know?
por: Cheng, Qinyuan, et al.
Publicado: (2024)
por: Cheng, Qinyuan, et al.
Publicado: (2024)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
por: Liu, Yilong, et al.
Publicado: (2026)
por: Liu, Yilong, et al.
Publicado: (2026)
Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets
por: Nel, Lukas
Publicado: (2025)
por: Nel, Lukas
Publicado: (2025)
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
por: Lee, Young-Suk, et al.
Publicado: (2026)
por: Lee, Young-Suk, et al.
Publicado: (2026)
FormGym: Doing Paperwork with Agents
por: Toles, Matthew, et al.
Publicado: (2025)
por: Toles, Matthew, et al.
Publicado: (2025)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
por: Song, Maojia, et al.
Publicado: (2025)
por: Song, Maojia, et al.
Publicado: (2025)
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
por: Jiang, Lei, et al.
Publicado: (2026)
por: Jiang, Lei, et al.
Publicado: (2026)
Human Resilience in the AI Era -- What Machines Can't Replace
por: Liu, Shaoshan, et al.
Publicado: (2025)
por: Liu, Shaoshan, et al.
Publicado: (2025)
Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant
por: Romanchuk, Oleg, et al.
Publicado: (2026)
por: Romanchuk, Oleg, et al.
Publicado: (2026)
K^2-Agent: Co-Evolving Know-What and Know-How for Hierarchical Mobile Device Control
por: Wu, Zhe, et al.
Publicado: (2026)
por: Wu, Zhe, et al.
Publicado: (2026)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
por: Peng, Jierui, et al.
Publicado: (2025)
por: Peng, Jierui, et al.
Publicado: (2025)
Budget-Aware Tool-Use Enables Effective Agent Scaling
por: Liu, Tengxiao, et al.
Publicado: (2025)
por: Liu, Tengxiao, et al.
Publicado: (2025)
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
por: Mannekote, Amogh, et al.
Publicado: (2025)
por: Mannekote, Amogh, et al.
Publicado: (2025)
I Know You Can't See Me: Dynamic Occlusion-Aware Safety Validation of Strategic Planners for Autonomous Vehicles Using Hypergames
por: Kahn, Maximilian, et al.
Publicado: (2021)
por: Kahn, Maximilian, et al.
Publicado: (2021)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
por: Wang, Ruipeng, et al.
Publicado: (2026)
por: Wang, Ruipeng, et al.
Publicado: (2026)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
por: Ma, Lu, et al.
Publicado: (2025)
por: Ma, Lu, et al.
Publicado: (2025)
Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning
por: Cui, Zhenyu, et al.
Publicado: (2026)
por: Cui, Zhenyu, et al.
Publicado: (2026)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
por: Ferrando, Javier, et al.
Publicado: (2024)
por: Ferrando, Javier, et al.
Publicado: (2024)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
por: Jung, Sungwoo, et al.
Publicado: (2026)
por: Jung, Sungwoo, et al.
Publicado: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
por: Guo, Zhengkang, et al.
Publicado: (2026)
por: Guo, Zhengkang, et al.
Publicado: (2026)
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
por: Hao, Chenjie, et al.
Publicado: (2026)
por: Hao, Chenjie, et al.
Publicado: (2026)
Agents Need Not Know Their Purpose
por: Garcia, Paulo
Publicado: (2024)
por: Garcia, Paulo
Publicado: (2024)
Do Retrieval Augmented Language Models Know When They Don't Know?
por: Zhou, Youchao, et al.
Publicado: (2025)
por: Zhou, Youchao, et al.
Publicado: (2025)
Why Do Multi-Agent LLM Systems Fail?
por: Cemri, Mert, et al.
Publicado: (2025)
por: Cemri, Mert, et al.
Publicado: (2025)
Ejemplares similares
-
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
por: Priyanshu, Aman, et al.
Publicado: (2026) -
What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding
por: Liu, Siyuan, et al.
Publicado: (2026) -
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
por: Shao, Jiaqi, et al.
Publicado: (2025) -
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
por: Cheng, Yize, et al.
Publicado: (2026) -
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
por: P, Vedanta S, et al.
Publicado: (2026)