The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Junlong, Zhao, Wenshuo, Zhao, Jian, Zeng, Weihao, Wu, Haoze, Wang, Xiaochen, Ge, Rui, Cao, Yuxuan, Huang, Yuzhen, Liu, Wei, Liu, Junteng, Su, Zhaochen, Guo, Yiyang, Zhou, Fan, Zhang, Lueyang, Michelini, Juan, Wang, Xingyao, Yue, Xiang, Zhou, Shuyan, Neubig, Graham, He, Junxian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910033354162176
author Li, Junlong
Zhao, Wenshuo
Zhao, Jian
Zeng, Weihao
Wu, Haoze
Wang, Xiaochen
Ge, Rui
Cao, Yuxuan
Huang, Yuzhen
Liu, Wei
Liu, Junteng
Su, Zhaochen
Guo, Yiyang
Zhou, Fan
Zhang, Lueyang
Michelini, Juan
Wang, Xingyao
Yue, Xiang
Zhou, Shuyan
Neubig, Graham
He, Junxian
author_facet Li, Junlong
Zhao, Wenshuo
Zhao, Jian
Zeng, Weihao
Wu, Haoze
Wang, Xiaochen
Ge, Rui
Cao, Yuxuan
Huang, Yuzhen
Liu, Wei
Liu, Junteng
Su, Zhaochen
Guo, Yiyang
Zhou, Fan
Zhang, Lueyang
Michelini, Juan
Wang, Xingyao
Yue, Xiang
Zhou, Shuyan
Neubig, Graham
He, Junxian
contents Real-world language agents must handle complex, multi-step workflows across diverse Apps. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database to detect anomalies and generate reports following an operating manual. However, existing language agent benchmarks often focus on narrow domains or simplified tasks that lack the diversity, realism, and long-horizon complexity required to evaluate agents' real-world performance. To address this gap, we introduce the Tool Decathlon (dubbed as Toolathlon), a benchmark for language agents offering diverse Apps and tools, realistic environment setup, and reliable execution-based evaluation. Toolathlon spans 32 software applications and 604 tools, ranging from everyday platforms such as Google Calendar and Notion to professional ones like WooCommerce, Kubernetes, and BigQuery. Most of the tools are based on a high-quality set of Model Context Protocol (MCP) servers that we may have revised or implemented ourselves. Unlike prior works, which primarily ensure functional realism but offer limited environment state diversity, we provide realistic initial environment states from real software, such as Canvas courses with dozens of students or real financial spreadsheets. This benchmark includes 108 manually sourced or crafted tasks in total, requiring interacting with multiple Apps over around 20 turns on average to complete. Each task is strictly verifiable through dedicated evaluation scripts. Comprehensive evaluation of SOTA models highlights their significant shortcomings: the best-performing model, Claude-4.5-Sonnet, achieves only a 38.6% success rate with 20.2 tool calling turns on average, while the top open-weights model DeepSeek-V3.2-Exp reaches 20.1%. We expect Toolathlon to drive the development of more capable language agents for real-world, long-horizon task execution.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Li, Junlong
Zhao, Wenshuo
Zhao, Jian
Zeng, Weihao
Wu, Haoze
Wang, Xiaochen
Ge, Rui
Cao, Yuxuan
Huang, Yuzhen
Liu, Wei
Liu, Junteng
Su, Zhaochen
Guo, Yiyang
Zhou, Fan
Zhang, Lueyang
Michelini, Juan
Wang, Xingyao
Yue, Xiang
Zhou, Shuyan
Neubig, Graham
He, Junxian
Computation and Language
Artificial Intelligence
Real-world language agents must handle complex, multi-step workflows across diverse Apps. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database to detect anomalies and generate reports following an operating manual. However, existing language agent benchmarks often focus on narrow domains or simplified tasks that lack the diversity, realism, and long-horizon complexity required to evaluate agents' real-world performance. To address this gap, we introduce the Tool Decathlon (dubbed as Toolathlon), a benchmark for language agents offering diverse Apps and tools, realistic environment setup, and reliable execution-based evaluation. Toolathlon spans 32 software applications and 604 tools, ranging from everyday platforms such as Google Calendar and Notion to professional ones like WooCommerce, Kubernetes, and BigQuery. Most of the tools are based on a high-quality set of Model Context Protocol (MCP) servers that we may have revised or implemented ourselves. Unlike prior works, which primarily ensure functional realism but offer limited environment state diversity, we provide realistic initial environment states from real software, such as Canvas courses with dozens of students or real financial spreadsheets. This benchmark includes 108 manually sourced or crafted tasks in total, requiring interacting with multiple Apps over around 20 turns on average to complete. Each task is strictly verifiable through dedicated evaluation scripts. Comprehensive evaluation of SOTA models highlights their significant shortcomings: the best-performing model, Claude-4.5-Sonnet, achieves only a 38.6% success rate with 20.2 tool calling turns on average, while the top open-weights model DeepSeek-V3.2-Exp reaches 20.1%. We expect Toolathlon to drive the development of more capable language agents for real-world, long-horizon task execution.
title The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.25726