GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jize, Liu, Xuanxuan, Li, Yining, Zhang, Songyang, Wang, Yijun, Shan, Zifei, Le, Xinyi, Chen, Cailian, Guan, Xinping, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GTA: A Benchmark for General Tool Agents
by: Wang, Jize, et al.
Published: (2024)
by: Wang, Jize, et al.
Published: (2024)
RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents
by: Wang, Jize, et al.
Published: (2026)
by: Wang, Jize, et al.
Published: (2026)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Private and Robust Distributed Nonconvex Optimization via Polynomial Approximation
by: He, Zhiyu, et al.
Published: (2021)
by: He, Zhiyu, et al.
Published: (2021)
Distributed Nonconvex Optimization: Gradient-free Iterations and $ε$-Globally Optimal Solution
by: He, Zhiyu, et al.
Published: (2020)
by: He, Zhiyu, et al.
Published: (2020)
Distributed Constraint-coupled Resource Allocation: Anytime Feasibility and Violation Robustness
by: Wu, Wenwen, et al.
Published: (2025)
by: Wu, Wenwen, et al.
Published: (2025)
Towards Verifiably Safe Tool Use for LLM Agents
by: Doshi, Aarya, et al.
Published: (2026)
by: Doshi, Aarya, et al.
Published: (2026)
Infinite-Horizon Optimal Wireless Control Over Shared State-Dependent Fading Channels for IIoT Systems
by: Wang, Shuling, et al.
Published: (2024)
by: Wang, Shuling, et al.
Published: (2024)
AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation
by: Wang, Siyu, et al.
Published: (2026)
by: Wang, Siyu, et al.
Published: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
by: Xu, Haoyuan, et al.
Published: (2026)
by: Xu, Haoyuan, et al.
Published: (2026)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
by: Luo, Jiaqi, et al.
Published: (2026)
by: Luo, Jiaqi, et al.
Published: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
Sample Complexity of Policy Gradient for Log-Growth Control
by: Pan, Qiuhua, et al.
Published: (2026)
by: Pan, Qiuhua, et al.
Published: (2026)
XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows
by: Wang, Xinru, et al.
Published: (2025)
by: Wang, Xinru, et al.
Published: (2025)
Procedural Environment Generation for Tool-Use Agents
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents
by: Yu, Bihui, et al.
Published: (2026)
by: Yu, Bihui, et al.
Published: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
OpenEP: Open-Ended Future Event Prediction
by: Guan, Yong, et al.
Published: (2024)
by: Guan, Yong, et al.
Published: (2024)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
The Tool Illusion: Rethinking Tool Use in Web Agents
by: Lou, Renze, et al.
Published: (2026)
by: Lou, Renze, et al.
Published: (2026)
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
by: Mohammadi, Bardia, et al.
Published: (2026)
by: Mohammadi, Bardia, et al.
Published: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
SAIL: Sample-Centric In-Context Learning for Document Information Extraction
by: Zhang, Jinyu, et al.
Published: (2024)
by: Zhang, Jinyu, et al.
Published: (2024)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
by: Zhao, Sijie, et al.
Published: (2026)
by: Zhao, Sijie, et al.
Published: (2026)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
by: Thaman, Kunvar
Published: (2026)
by: Thaman, Kunvar
Published: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
GTA: Generative Traffic Agents for Simulating Realistic Mobility Behavior
by: Lämmer, Simon, et al.
Published: (2026)
by: Lämmer, Simon, et al.
Published: (2026)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
CCTU: A Benchmark for Tool Use under Complex Constraints
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
LLM Agents Beyond Utility: An Open-Ended Perspective
by: Nachkov, Asen, et al.
Published: (2025)
by: Nachkov, Asen, et al.
Published: (2025)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
by: Huang, Yue, et al.
Published: (2023)
by: Huang, Yue, et al.
Published: (2023)
BulkSeq Workflow Tools
by: Manzone Rodriguez, Clarisa, et al.
Published: (2026)
by: Manzone Rodriguez, Clarisa, et al.
Published: (2026)
Multi-Agent Craftax: Benchmarking Open-Ended Multi-Agent Reinforcement Learning at the Hyperscale
by: Omari, Bassel Al, et al.
Published: (2025)
by: Omari, Bassel Al, et al.
Published: (2025)
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
by: Guo, Ruocheng, et al.
Published: (2026)
by: Guo, Ruocheng, et al.
Published: (2026)
Similar Items
-
GTA: A Benchmark for General Tool Agents
by: Wang, Jize, et al.
Published: (2024) -
RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents
by: Wang, Jize, et al.
Published: (2026) -
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025) -
Private and Robust Distributed Nonconvex Optimization via Polynomial Approximation
by: He, Zhiyu, et al.
Published: (2021) -
Distributed Nonconvex Optimization: Gradient-free Iterations and $ε$-Globally Optimal Solution
by: He, Zhiyu, et al.
Published: (2020)