Tools Fail: Detecting Silent Errors in Faulty Tools
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Jimin, Min, So Yeon, Chang, Yingshan, Bisk, Yonatan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models Need Inductive Biases to Count Inductively
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
Self-Regulation and Requesting Interventions
von: Min, So Yeon, et al.
Veröffentlicht: (2025)
von: Min, So Yeon, et al.
Veröffentlicht: (2025)
Current Agents Fail to Leverage World Model as Tool for Foresight
von: Qian, Cheng, et al.
Veröffentlicht: (2026)
von: Qian, Cheng, et al.
Veröffentlicht: (2026)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Learning Model Successors
von: Chang, Yingshan, et al.
Veröffentlicht: (2025)
von: Chang, Yingshan, et al.
Veröffentlicht: (2025)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
Tool Unlearning for Tool-Augmented LLMs
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
Classification of User Reports for Detection of Faulty Computer Components using NLP Models: A Case Study
von: Silva, Maria de Lourdes M., et al.
Veröffentlicht: (2025)
von: Silva, Maria de Lourdes M., et al.
Veröffentlicht: (2025)
ToolRL: Reward is All Tool Learning Needs
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
Tool Learning with Foundation Models
von: Qin, Yujia, et al.
Veröffentlicht: (2023)
von: Qin, Yujia, et al.
Veröffentlicht: (2023)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
von: Zhu, Dongsheng, et al.
Veröffentlicht: (2025)
von: Zhu, Dongsheng, et al.
Veröffentlicht: (2025)
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
von: Zeng, Yirong, et al.
Veröffentlicht: (2025)
von: Zeng, Yirong, et al.
Veröffentlicht: (2025)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
ChemAmp: Amplified Chemistry Tools via Composable Agents
von: Li, Zhucong, et al.
Veröffentlicht: (2025)
von: Li, Zhucong, et al.
Veröffentlicht: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play
von: Fang, Wei, et al.
Veröffentlicht: (2025)
von: Fang, Wei, et al.
Veröffentlicht: (2025)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
Large Language Models as Tool Makers
von: Cai, Tianle, et al.
Veröffentlicht: (2023)
von: Cai, Tianle, et al.
Veröffentlicht: (2023)
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
CNSight: Evaluation of Clinical Note Segmentation Tools
von: Surana, Risha, et al.
Veröffentlicht: (2025)
von: Surana, Risha, et al.
Veröffentlicht: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
ToolACE: Winning the Points of LLM Function Calling
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
SMART: Self-Aware Agent for Tool Overuse Mitigation
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
von: Su, Hongjin, et al.
Veröffentlicht: (2025)
von: Su, Hongjin, et al.
Veröffentlicht: (2025)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
von: Gowda, Thamme, et al.
Veröffentlicht: (2021)
von: Gowda, Thamme, et al.
Veröffentlicht: (2021)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language Models Need Inductive Biases to Count Inductively
von: Chang, Yingshan, et al.
Veröffentlicht: (2024) -
Self-Regulation and Requesting Interventions
von: Min, So Yeon, et al.
Veröffentlicht: (2025) -
Current Agents Fail to Leverage World Model as Tool for Foresight
von: Qian, Cheng, et al.
Veröffentlicht: (2026) -
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025) -
Learning Model Successors
von: Chang, Yingshan, et al.
Veröffentlicht: (2025)