ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Hyeonje, Lee, Jeongsoo, Lee, Hyojun, Lee, Jay-Yoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
von: Ni, Xinyi, et al.
Veröffentlicht: (2025)
von: Ni, Xinyi, et al.
Veröffentlicht: (2025)
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
von: Liu, Marianne Menglin, et al.
Veröffentlicht: (2025)
von: Liu, Marianne Menglin, et al.
Veröffentlicht: (2025)
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
von: Wang, Shuaiqi, et al.
Veröffentlicht: (2026)
von: Wang, Shuaiqi, et al.
Veröffentlicht: (2026)
Tool Calling is Linearly Readable and Steerable in Language Models
von: Wu, Zekun, et al.
Veröffentlicht: (2026)
von: Wu, Zekun, et al.
Veröffentlicht: (2026)
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
Benchmarking Failures in Tool-Augmented Language Models
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
von: Li, Xu, et al.
Veröffentlicht: (2026)
von: Li, Xu, et al.
Veröffentlicht: (2026)
Good Tools are Half the Work: Tool Usage in Deep Learning Projects
von: Panourgia, Evangelia, et al.
Veröffentlicht: (2023)
von: Panourgia, Evangelia, et al.
Veröffentlicht: (2023)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
von: Chen, Shiqi, et al.
Veröffentlicht: (2026)
von: Chen, Shiqi, et al.
Veröffentlicht: (2026)
Planning-Aware Code Infilling via Horizon-Length Prediction
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
AutoGen Studio: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems
von: Dibia, Victor, et al.
Veröffentlicht: (2024)
von: Dibia, Victor, et al.
Veröffentlicht: (2024)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
von: LeVine, Will, et al.
Veröffentlicht: (2026)
von: LeVine, Will, et al.
Veröffentlicht: (2026)
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration
von: Shen, Yuanhao, et al.
Veröffentlicht: (2024)
von: Shen, Yuanhao, et al.
Veröffentlicht: (2024)
Sanskrit Knowledge-based Systems: Annotation and Computational Tools
von: Terdalkar, Hrishikesh
Veröffentlicht: (2024)
von: Terdalkar, Hrishikesh
Veröffentlicht: (2024)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
von: Castellani, Tommaso, et al.
Veröffentlicht: (2025)
von: Castellani, Tommaso, et al.
Veröffentlicht: (2025)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
von: Lee, Seongmin, et al.
Veröffentlicht: (2025)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
RepoQA: Evaluating Long Context Code Understanding
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
Lessons from the Use of Natural Language Inference (NLI) in Requirements Engineering Tasks
von: Fazelnia, Mohamad, et al.
Veröffentlicht: (2024)
von: Fazelnia, Mohamad, et al.
Veröffentlicht: (2024)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
von: Golubev, Alexander, et al.
Veröffentlicht: (2025)
von: Golubev, Alexander, et al.
Veröffentlicht: (2025)
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
von: Li, Yufei, et al.
Veröffentlicht: (2024)
von: Li, Yufei, et al.
Veröffentlicht: (2024)
Towards Continuous Assurance Case Creation for ADS with the Evidential Tool Bus
von: Sorokin, Lev, et al.
Veröffentlicht: (2024)
von: Sorokin, Lev, et al.
Veröffentlicht: (2024)
Narrowing the Gap: Supervised Fine-Tuning of Open-Source LLMs as a Viable Alternative to Proprietary Models for Pedagogical Tools
von: Solano, Lorenzo Lee, et al.
Veröffentlicht: (2025)
von: Solano, Lorenzo Lee, et al.
Veröffentlicht: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
von: Wang, Qihao, et al.
Veröffentlicht: (2026)
von: Wang, Qihao, et al.
Veröffentlicht: (2026)
RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models
von: Wang, Zefan, et al.
Veröffentlicht: (2023)
von: Wang, Zefan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023) -
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
von: Lu, Yifei, et al.
Veröffentlicht: (2025) -
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
von: Iskander, Shadi, et al.
Veröffentlicht: (2024) -
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
von: Ni, Xinyi, et al.
Veröffentlicht: (2025) -
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
von: Ding, Peng, et al.
Veröffentlicht: (2025)