CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Shiting, Fang, Zhen, Chen, Zehui, Yuan, Siyu, Ye, Junjie, Zeng, Yu, Chen, Lin, Mao, Qi, Zhao, Feng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913896479064064
author Huang, Shiting
Fang, Zhen
Chen, Zehui
Yuan, Siyu
Ye, Junjie
Zeng, Yu
Chen, Lin
Mao, Qi
Zhao, Feng
author_facet Huang, Shiting
Fang, Zhen
Chen, Zehui
Yuan, Siyu
Ye, Junjie
Zeng, Yu
Chen, Lin
Mao, Qi
Zhao, Feng
contents The ability of large language models (LLMs) to utilize external tools has enabled them to tackle an increasingly diverse range of tasks. However, as the tasks become more complex and long-horizon, the intricate tool utilization process may trigger various unexpected errors. Therefore, how to effectively handle such errors, including identifying, diagnosing, and recovering from them, has emerged as a key research direction for advancing tool learning. In this work, we first extensively analyze the types of errors encountered during the function-calling process on several competitive tool evaluation benchmarks. Based on it, we introduce CRITICTOOL, a comprehensive critique evaluation benchmark specialized for tool learning. Building upon a novel evolutionary strategy for dataset construction, CRITICTOOL holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. We conduct extensive experiments on CRITICTOOL, and validate the generalization and effectiveness of our constructed benchmark strategy. We also provide an in-depth analysis of the tool reflection ability on various LLMs, offering a new perspective on the field of tool learning in LLMs. The code is available at \href{https://github.com/Shellorley0513/CriticTool}{https://github.com/Shellorley0513/CriticTool}.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
Huang, Shiting
Fang, Zhen
Chen, Zehui
Yuan, Siyu
Ye, Junjie
Zeng, Yu
Chen, Lin
Mao, Qi
Zhao, Feng
Software Engineering
Computation and Language
The ability of large language models (LLMs) to utilize external tools has enabled them to tackle an increasingly diverse range of tasks. However, as the tasks become more complex and long-horizon, the intricate tool utilization process may trigger various unexpected errors. Therefore, how to effectively handle such errors, including identifying, diagnosing, and recovering from them, has emerged as a key research direction for advancing tool learning. In this work, we first extensively analyze the types of errors encountered during the function-calling process on several competitive tool evaluation benchmarks. Based on it, we introduce CRITICTOOL, a comprehensive critique evaluation benchmark specialized for tool learning. Building upon a novel evolutionary strategy for dataset construction, CRITICTOOL holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. We conduct extensive experiments on CRITICTOOL, and validate the generalization and effectiveness of our constructed benchmark strategy. We also provide an in-depth analysis of the tool reflection ability on various LLMs, offering a new perspective on the field of tool learning in LLMs. The code is available at \href{https://github.com/Shellorley0513/CriticTool}{https://github.com/Shellorley0513/CriticTool}.
title CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2506.13977