ZEBRAARENA: A Diagnostic Simulation Environment for Studying Reasoning-Action Coupling in Tool-Augmented LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Wanjia, Schmidt, Ludwig, Zou, James, Balachandran, Vidhisha, Chen, Lingjiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
by: Zhao, Wanjia, et al.
Published: (2025)
by: Zhao, Wanjia, et al.
Published: (2025)
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
by: Vilas, Martina G., et al.
Published: (2025)
by: Vilas, Martina G., et al.
Published: (2025)
Reasoning Up the Instruction Ladder for Controllable Language Models
by: Zheng, Zishuo, et al.
Published: (2025)
by: Zheng, Zishuo, et al.
Published: (2025)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2024)
by: Li, Shuyue Stella, et al.
Published: (2024)
Knowledge Crosswords: Geometric Knowledge Reasoning with Large Language Models
by: Ding, Wenxuan, et al.
Published: (2023)
by: Ding, Wenxuan, et al.
Published: (2023)
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
by: Balachandran, Vidhisha, et al.
Published: (2025)
by: Balachandran, Vidhisha, et al.
Published: (2025)
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
by: Du, Yongkang, et al.
Published: (2026)
by: Du, Yongkang, et al.
Published: (2026)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters
by: Zhao, Wanjia, et al.
Published: (2025)
by: Zhao, Wanjia, et al.
Published: (2025)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
by: Wu, Wenxun, et al.
Published: (2025)
by: Wu, Wenxun, et al.
Published: (2025)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations
by: Song, Zhao, et al.
Published: (2025)
by: Song, Zhao, et al.
Published: (2025)
Topology of Reasoning: Retrieved Cell Complex-Augmented Generation for Textual Graph Question Answering
by: Zhao, Sen, et al.
Published: (2026)
by: Zhao, Sen, et al.
Published: (2026)
Investigating Tool-Memory Conflicts in Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2026)
by: Cheng, Jiali, et al.
Published: (2026)
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning
by: Li, Yize, et al.
Published: (2026)
by: Li, Yize, et al.
Published: (2026)
MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs
by: Hsu, Hsin-Ling, et al.
Published: (2026)
by: Hsu, Hsin-Ling, et al.
Published: (2026)
Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
by: Wei, Xiaolong, et al.
Published: (2025)
by: Wei, Xiaolong, et al.
Published: (2025)
Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning
by: Pan, Zhikai, et al.
Published: (2026)
by: Pan, Zhikai, et al.
Published: (2026)
CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
by: Shi, Dachuan, et al.
Published: (2026)
by: Shi, Dachuan, et al.
Published: (2026)
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026)
by: Gao, Xingjie, et al.
Published: (2026)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration
by: Deng, Liyuan, et al.
Published: (2026)
by: Deng, Liyuan, et al.
Published: (2026)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025)
by: Zhao, Yufeng, et al.
Published: (2025)
Simulating Environments with Reasoning Models for Agent Training
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
by: Li, Kuan, et al.
Published: (2026)
by: Li, Kuan, et al.
Published: (2026)
Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents
by: Enabe, Paulo Akira F.
Published: (2026)
by: Enabe, Paulo Akira F.
Published: (2026)
Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning
by: Jiang, Chuang, et al.
Published: (2025)
by: Jiang, Chuang, et al.
Published: (2025)
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
by: Yang, Chen, et al.
Published: (2025)
by: Yang, Chen, et al.
Published: (2025)
Guideline Forest: Retrieval-Augmented Reasoning with Branching Experience-Induced Guidelines
by: Chen, Jiaxiang, et al.
Published: (2025)
by: Chen, Jiaxiang, et al.
Published: (2025)
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
by: Zhang, Situo, et al.
Published: (2026)
by: Zhang, Situo, et al.
Published: (2026)
Reasoning and Sampling-Augmented MCQ Difficulty Prediction via LLMs
by: Feng, Wanyong, et al.
Published: (2025)
by: Feng, Wanyong, et al.
Published: (2025)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
by: Zeng, Yirong, et al.
Published: (2025)
by: Zeng, Yirong, et al.
Published: (2025)
From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting
by: Chen, Nuo, et al.
Published: (2023)
by: Chen, Nuo, et al.
Published: (2023)
AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning
by: Tong, Siqian, et al.
Published: (2026)
by: Tong, Siqian, et al.
Published: (2026)
Similar Items
-
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
by: Zhao, Wanjia, et al.
Published: (2025) -
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
by: Vilas, Martina G., et al.
Published: (2025) -
Reasoning Up the Instruction Ladder for Controllable Language Models
by: Zheng, Zishuo, et al.
Published: (2025) -
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2024) -
Knowledge Crosswords: Geometric Knowledge Reasoning with Large Language Models
by: Ding, Wenxuan, et al.
Published: (2023)