ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nath, Vaskar, Raja, Pranav, Yoon, Claire, Hendryx, Sean |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
von: Nath, Vaskar, et al.
Veröffentlicht: (2025)
Learning Goal-Conditioned Representations for Language Reward Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2024)
von: Nath, Vaskar, et al.
Veröffentlicht: (2024)
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
von: Jin, Alwin, et al.
Veröffentlicht: (2025)
von: Jin, Alwin, et al.
Veröffentlicht: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
Planning In Natural Language Improves LLM Search For Code Generation
von: Wang, Evan, et al.
Veröffentlicht: (2024)
von: Wang, Evan, et al.
Veröffentlicht: (2024)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
von: Choi, Hyeonje, et al.
Veröffentlicht: (2026)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
The Tool Illusion: Rethinking Tool Use in Web Agents
von: Lou, Renze, et al.
Veröffentlicht: (2026)
von: Lou, Renze, et al.
Veröffentlicht: (2026)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
von: LeVine, Will, et al.
Veröffentlicht: (2023)
von: LeVine, Will, et al.
Veröffentlicht: (2023)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
von: Xu, Ruotao, et al.
Veröffentlicht: (2026)
von: Xu, Ruotao, et al.
Veröffentlicht: (2026)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
von: Zhang, Yindong, et al.
Veröffentlicht: (2026)
von: Zhang, Yindong, et al.
Veröffentlicht: (2026)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2025)
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2025)
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
von: Geng, Xinyu, et al.
Veröffentlicht: (2026)
von: Geng, Xinyu, et al.
Veröffentlicht: (2026)
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
START: Self-taught Reasoner with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
von: Wu, Mengsong, et al.
Veröffentlicht: (2024)
von: Wu, Mengsong, et al.
Veröffentlicht: (2024)
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval
von: Chen, Yanfei, et al.
Veröffentlicht: (2024)
von: Chen, Yanfei, et al.
Veröffentlicht: (2024)
MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching
von: Qu, Changle, et al.
Veröffentlicht: (2026)
von: Qu, Changle, et al.
Veröffentlicht: (2026)
MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications
von: Vinayagame, Vishnou, et al.
Veröffentlicht: (2024)
von: Vinayagame, Vishnou, et al.
Veröffentlicht: (2024)
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
von: Wang, Pei, et al.
Veröffentlicht: (2024)
von: Wang, Pei, et al.
Veröffentlicht: (2024)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
von: Wang, Li, et al.
Veröffentlicht: (2026)
von: Wang, Li, et al.
Veröffentlicht: (2026)
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
von: Ma, Qinghe, et al.
Veröffentlicht: (2026)
Efficient Tool Use with Chain-of-Abstraction Reasoning
von: Gao, Silin, et al.
Veröffentlicht: (2024)
von: Gao, Silin, et al.
Veröffentlicht: (2024)
MatTools: Benchmarking Large Language Models for Materials Science Tools
von: Liu, Siyu, et al.
Veröffentlicht: (2025)
von: Liu, Siyu, et al.
Veröffentlicht: (2025)
GTA: A Benchmark for General Tool Agents
von: Wang, Jize, et al.
Veröffentlicht: (2024)
von: Wang, Jize, et al.
Veröffentlicht: (2024)
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
von: Guo, Zhicheng, et al.
Veröffentlicht: (2024)
von: Guo, Zhicheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024) -
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2025) -
Learning Goal-Conditioned Representations for Language Reward Models
von: Nath, Vaskar, et al.
Veröffentlicht: (2024) -
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
von: Jin, Alwin, et al.
Veröffentlicht: (2025) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)