ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kokane, Shirley, Zhu, Ming, Awalgaonkar, Tulika, Zhang, Jianguo, Hoang, Thai, Prabhakar, Akshara, Liu, Zuxin, Lan, Tian, Yang, Liangwei, Tan, Juntao, Murthy, Rithesh, Yao, Weiran, Liu, Zhiwei, Niebles, Juan Carlos, Wang, Huan, Heinecke, Shelby, Xiong, Caiming, Savarese, Silivo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915359149260800
author Kokane, Shirley
Zhu, Ming
Awalgaonkar, Tulika
Zhang, Jianguo
Hoang, Thai
Prabhakar, Akshara
Liu, Zuxin
Lan, Tian
Yang, Liangwei
Tan, Juntao
Murthy, Rithesh
Yao, Weiran
Liu, Zhiwei
Niebles, Juan Carlos
Wang, Huan
Heinecke, Shelby
Xiong, Caiming
Savarese, Silivo
author_facet Kokane, Shirley
Zhu, Ming
Awalgaonkar, Tulika
Zhang, Jianguo
Hoang, Thai
Prabhakar, Akshara
Liu, Zuxin
Lan, Tian
Yang, Liangwei
Tan, Juntao
Murthy, Rithesh
Yao, Weiran
Liu, Zhiwei
Niebles, Juan Carlos
Wang, Huan
Heinecke, Shelby
Xiong, Caiming
Savarese, Silivo
contents Evaluating Large Language Models (LLMs) is one of the most critical aspects of building a performant compound AI system. Since the output from LLMs propagate to downstream steps, identifying LLM errors is crucial to system performance. A common task for LLMs in AI systems is tool use. While there are several benchmark environments for evaluating LLMs on this task, they typically only give a success rate without any explanation of the failure cases. To solve this problem, we introduce TOOLSCAN, a new benchmark to identify error patterns in LLM output on tool-use tasks. Our benchmark data set comprises of queries from diverse environments that can be used to test for the presence of seven newly characterized error patterns. Using TOOLSCAN, we show that even the most prominent LLMs exhibit these error patterns in their outputs. Researchers can use these insights from TOOLSCAN to guide their error mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
Kokane, Shirley
Zhu, Ming
Awalgaonkar, Tulika
Zhang, Jianguo
Hoang, Thai
Prabhakar, Akshara
Liu, Zuxin
Lan, Tian
Yang, Liangwei
Tan, Juntao
Murthy, Rithesh
Yao, Weiran
Liu, Zhiwei
Niebles, Juan Carlos
Wang, Huan
Heinecke, Shelby
Xiong, Caiming
Savarese, Silivo
Software Engineering
Artificial Intelligence
Evaluating Large Language Models (LLMs) is one of the most critical aspects of building a performant compound AI system. Since the output from LLMs propagate to downstream steps, identifying LLM errors is crucial to system performance. A common task for LLMs in AI systems is tool use. While there are several benchmark environments for evaluating LLMs on this task, they typically only give a success rate without any explanation of the failure cases. To solve this problem, we introduce TOOLSCAN, a new benchmark to identify error patterns in LLM output on tool-use tasks. Our benchmark data set comprises of queries from diverse environments that can be used to test for the presence of seven newly characterized error patterns. Using TOOLSCAN, we show that even the most prominent LLMs exhibit these error patterns in their outputs. Researchers can use these insights from TOOLSCAN to guide their error mitigation strategies.
title ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2411.13547