ToLeaP: Rethinking Development of Tool Learning with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Haotian, Song, Zijun, Niu, Boye, Zhang, Ke, Ou, Litu, Lu, Yaxi, Zhang, Zhong, Cong, Xin, Lin, Yankai, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915291595800576
author Chen, Haotian
Song, Zijun
Niu, Boye
Zhang, Ke
Ou, Litu
Lu, Yaxi
Zhang, Zhong
Cong, Xin
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
author_facet Chen, Haotian
Song, Zijun
Niu, Boye
Zhang, Ke
Ou, Litu
Lu, Yaxi
Zhang, Zhong
Cong, Xin
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
contents Tool learning, which enables large language models (LLMs) to utilize external tools effectively, has garnered increasing attention for its potential to revolutionize productivity across industries. Despite rapid development in tool learning, key challenges and opportunities remain understudied, limiting deeper insights and future advancements. In this paper, we investigate the tool learning ability of 41 prevalent LLMs by reproducing 33 benchmarks and enabling one-click evaluation for seven of them, forming a Tool Learning Platform named ToLeaP. We also collect 21 out of 33 potential training datasets to facilitate future exploration. After analyzing over 3,000 bad cases of 41 LLMs based on ToLeaP, we identify four main critical challenges: (1) benchmark limitations induce both the neglect and lack of (2) autonomous learning, (3) generalization, and (4) long-horizon task-solving capabilities of LLMs. To aid future advancements, we take a step further toward exploring potential directions, namely (1) real-world benchmark construction, (2) compatibility-aware autonomous learning, (3) rationale learning by thinking, and (4) identifying and recalling key clues. The preliminary experiments demonstrate their effectiveness, highlighting the need for further research and exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ToLeaP: Rethinking Development of Tool Learning with Large Language Models
Chen, Haotian
Song, Zijun
Niu, Boye
Zhang, Ke
Ou, Litu
Lu, Yaxi
Zhang, Zhong
Cong, Xin
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
Artificial Intelligence
Tool learning, which enables large language models (LLMs) to utilize external tools effectively, has garnered increasing attention for its potential to revolutionize productivity across industries. Despite rapid development in tool learning, key challenges and opportunities remain understudied, limiting deeper insights and future advancements. In this paper, we investigate the tool learning ability of 41 prevalent LLMs by reproducing 33 benchmarks and enabling one-click evaluation for seven of them, forming a Tool Learning Platform named ToLeaP. We also collect 21 out of 33 potential training datasets to facilitate future exploration. After analyzing over 3,000 bad cases of 41 LLMs based on ToLeaP, we identify four main critical challenges: (1) benchmark limitations induce both the neglect and lack of (2) autonomous learning, (3) generalization, and (4) long-horizon task-solving capabilities of LLMs. To aid future advancements, we take a step further toward exploring potential directions, namely (1) real-world benchmark construction, (2) compatibility-aware autonomous learning, (3) rationale learning by thinking, and (4) identifying and recalling key clues. The preliminary experiments demonstrate their effectiveness, highlighting the need for further research and exploration.
title ToLeaP: Rethinking Development of Tool Learning with Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2505.11833