MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Pei, Wu, Yanan, Wang, Zekun, Liu, Jiaheng, Song, Xiaoshuai, Peng, Zhongyuan, Deng, Ken, Zhang, Chenchen, Wang, Jiakai, Peng, Junran, Zhang, Ge, Guo, Hangyu, Zhang, Zhaoxiang, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909350428147712
author Wang, Pei
Wu, Yanan
Wang, Zekun
Liu, Jiaheng
Song, Xiaoshuai
Peng, Zhongyuan
Deng, Ken
Zhang, Chenchen
Wang, Jiakai
Peng, Junran
Zhang, Ge
Guo, Hangyu
Zhang, Zhaoxiang
Su, Wenbo
Zheng, Bo
author_facet Wang, Pei
Wu, Yanan
Wang, Zekun
Liu, Jiaheng
Song, Xiaoshuai
Peng, Zhongyuan
Deng, Ken
Zhang, Chenchen
Wang, Jiakai
Peng, Junran
Zhang, Ge
Guo, Hangyu
Zhang, Zhaoxiang
Su, Wenbo
Zheng, Bo
contents Large Language Models (LLMs) have displayed massive improvements in reasoning and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing datasets have the following limitations: (1). Insufficient evaluation scenarios (e.g., only cover limited tool-use scenes). (2). Extensive evaluation costs (e.g., GPT API costs). To address these limitations, in this work, we propose a multi-granularity tool-use benchmark for large language models called MTU-Bench. For the "multi-granularity" property, our MTU-Bench covers five tool usage scenes (i.e., single-turn and single-tool, single-turn and multiple-tool, multiple-turn and single-tool, multiple-turn and multiple-tool, and out-of-distribution tasks). Besides, all evaluation metrics of our MTU-Bench are based on the prediction results and the ground truth without using any GPT or human evaluation metrics. Moreover, our MTU-Bench is collected by transforming existing high-quality datasets to simulate real-world tool usage scenarios, and we also propose an instruction dataset called MTU-Instruct data to enhance the tool-use abilities of existing LLMs. Comprehensive experimental results demonstrate the effectiveness of our MTU-Bench. Code and data will be released at https: //github.com/MTU-Bench-Team/MTU-Bench.git.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11710
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
Wang, Pei
Wu, Yanan
Wang, Zekun
Liu, Jiaheng
Song, Xiaoshuai
Peng, Zhongyuan
Deng, Ken
Zhang, Chenchen
Wang, Jiakai
Peng, Junran
Zhang, Ge
Guo, Hangyu
Zhang, Zhaoxiang
Su, Wenbo
Zheng, Bo
Computation and Language
Large Language Models (LLMs) have displayed massive improvements in reasoning and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing datasets have the following limitations: (1). Insufficient evaluation scenarios (e.g., only cover limited tool-use scenes). (2). Extensive evaluation costs (e.g., GPT API costs). To address these limitations, in this work, we propose a multi-granularity tool-use benchmark for large language models called MTU-Bench. For the "multi-granularity" property, our MTU-Bench covers five tool usage scenes (i.e., single-turn and single-tool, single-turn and multiple-tool, multiple-turn and single-tool, multiple-turn and multiple-tool, and out-of-distribution tasks). Besides, all evaluation metrics of our MTU-Bench are based on the prediction results and the ground truth without using any GPT or human evaluation metrics. Moreover, our MTU-Bench is collected by transforming existing high-quality datasets to simulate real-world tool usage scenarios, and we also propose an instruction dataset called MTU-Instruct data to enhance the tool-use abilities of existing LLMs. Comprehensive experimental results demonstrate the effectiveness of our MTU-Bench. Code and data will be released at https: //github.com/MTU-Bench-Team/MTU-Bench.git.
title MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.11710