The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junlong, Zhao, Wenshuo, Zhao, Jian, Zeng, Weihao, Wu, Haoze, Wang, Xiaochen, Ge, Rui, Cao, Yuxuan, Huang, Yuzhen, Liu, Wei, Liu, Junteng, Su, Zhaochen, Guo, Yiyang, Zhou, Fan, Zhang, Lueyang, Michelini, Juan, Wang, Xingyao, Yue, Xiang, Zhou, Shuyan, Neubig, Graham, He, Junxian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
by: Wu, Haoze, et al.
Published: (2025)
by: Wu, Haoze, et al.
Published: (2025)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026)
by: Su, Zhaochen, et al.
Published: (2026)
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
by: Zeng, Weihao, et al.
Published: (2026)
by: Zeng, Weihao, et al.
Published: (2026)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
by: Wang, Xingyao, et al.
Published: (2026)
by: Wang, Xingyao, et al.
Published: (2026)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
by: Wang, Xingyao, et al.
Published: (2025)
by: Wang, Xingyao, et al.
Published: (2025)
How can we assess human-agent interactions? Case studies in software agent design
by: Chen, Valerie, et al.
Published: (2025)
by: Chen, Valerie, et al.
Published: (2025)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025)
by: Soni, Aditya Bharat, et al.
Published: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
by: Zhou, Shuyan, et al.
Published: (2023)
by: Zhou, Shuyan, et al.
Published: (2023)
On the Universal Truthfulness Hyperplane Inside LLMs
by: Liu, Junteng, et al.
Published: (2024)
by: Liu, Junteng, et al.
Published: (2024)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
by: Huang, Yuzhen, et al.
Published: (2025)
by: Huang, Yuzhen, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Quantum Optimization Benchmarking Library - The Intractable Decathlon
by: Koch, Thorsten, et al.
Published: (2025)
by: Koch, Thorsten, et al.
Published: (2025)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
by: Zeng, Weihao, et al.
Published: (2025)
by: Zeng, Weihao, et al.
Published: (2025)
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
Diving into Self-Evolving Training for Multimodal Reasoning
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
Layer-Dependent Quantum Anomalous Hall Effect in Rhombohedral Graphene
by: Liu, Zhaochen, et al.
Published: (2024)
by: Liu, Zhaochen, et al.
Published: (2024)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
by: Liu, Emmy, et al.
Published: (2024)
by: Liu, Emmy, et al.
Published: (2024)
Midtraining Bridges Pretraining and Posttraining Distributions
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
by: Zhou, Ruochen, et al.
Published: (2025)
by: Zhou, Ruochen, et al.
Published: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
by: Wang, Zhiruo, et al.
Published: (2024)
by: Wang, Zhiruo, et al.
Published: (2024)
Go-Browse: Training Web Agents with Structured Exploration
by: Gandhi, Apurva, et al.
Published: (2025)
by: Gandhi, Apurva, et al.
Published: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
by: Tjuatja, Lindia, et al.
Published: (2025)
by: Tjuatja, Lindia, et al.
Published: (2025)
Effective Strategies for Asynchronous Software Engineering Agents
by: Geng, Jiayi, et al.
Published: (2026)
by: Geng, Jiayi, et al.
Published: (2026)
Divergences between Language Models and Human Brains
by: Zhou, Yuchen, et al.
Published: (2023)
by: Zhou, Yuchen, et al.
Published: (2023)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
by: Xie, Yiqing, et al.
Published: (2026)
by: Xie, Yiqing, et al.
Published: (2026)
NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents
by: Song, Yang, et al.
Published: (2026)
by: Song, Yang, et al.
Published: (2026)
SWE-RM: Execution-free Feedback For Software Engineering Agents
by: Shum, KaShun, et al.
Published: (2025)
by: Shum, KaShun, et al.
Published: (2025)
PACC: Protocol-Aware Cross-Layer Compression for Compact Network Traffic Representation
by: Guo, Zhaochen, et al.
Published: (2026)
by: Guo, Zhaochen, et al.
Published: (2026)
Knowledge Distillation Must Account for What It Loses
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
LLM Reasoning Is Latent, Not the Chain of Thought
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
LLMs Should Not Yet Be Credited with Decision Explanation
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
by: Ou, Tianyue, et al.
Published: (2024)
by: Ou, Tianyue, et al.
Published: (2024)
Similar Items
-
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
by: Wu, Haoze, et al.
Published: (2025) -
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026) -
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
by: Zeng, Weihao, et al.
Published: (2026) -
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024) -
A Rubric-Supervised Critic from Sparse Real-World Outcomes
by: Wang, Xingyao, et al.
Published: (2026)