When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Xiaolin, Yuan, Aojie, Luo, Zheng, Ling, Zipeng, Pan, Xixiao, Gao, Yicheng, Zhang, Haiyue, Li, Jiate, Jiang, Shuli, Wang, Prince Zizhuang, Zhu, Zixuan, Liu, Jinbo, Rossi, Ryan A., Wei, Hua, Hu, Xiyang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913116739076096
author Zhou, Xiaolin
Yuan, Aojie
Luo, Zheng
Ling, Zipeng
Pan, Xixiao
Gao, Yicheng
Zhang, Haiyue
Li, Jiate
Jiang, Shuli
Wang, Prince Zizhuang
Zhu, Zixuan
Liu, Jinbo
Rossi, Ryan A.
Wei, Hua
Hu, Xiyang
author_facet Zhou, Xiaolin
Yuan, Aojie
Luo, Zheng
Ling, Zipeng
Pan, Xixiao
Gao, Yicheng
Zhang, Haiyue
Li, Jiate
Jiang, Shuli
Wang, Prince Zizhuang
Zhu, Zixuan
Liu, Jinbo
Rossi, Ryan A.
Wei, Hua
Hu, Xiyang
contents Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user typos propagate into hallucinated tool names, a misconfigured request timeout can stall an agent indefinitely, and duplicate tool names across servers can freeze an SDK. We study these failures as a sim-to-real gap in the tool-use partially observable Markov decision process (POMDP), where deployment noise enters through the observation, action space, reward-relevant metadata, or transition dynamics. We introduce RobustBench-TC, a benchmark with 22 perturbation types organized by these four POMDP components, each grounded in a verified GitHub issue or documented tool-calling failure. Across 21 models from 1.5B to 32B parameters (including the closed-source o4-mini), the robustness profile is sharply uneven: observation perturbations reduce accuracy by less than 5%, while reward-relevant and transition perturbations reduce accuracy by roughly 40% and 30%, respectively; scale alone does not close these gaps. We then propose ToolRL-DR, a domain-randomization reinforcement learning (RL) recipe that trains a tool-use agent on perturbation-augmented trajectories spanning the three statically encodable POMDP components. On a 3B backbone, ToolRL-DR-Full retains roughly three-quarters of clean accuracy and reaches an aggregate perturbed accuracy comparable to open-source 14B function-calling baselines while substantially narrowing the gap to o4-mini. It closes approximately 27% of the Transition gap despite never seeing transition perturbations in training, suggesting that RL on adversarial static tool-use inputs induces a more persistent retry policy that transfers to unseen runtime failures. The dataset, code and benchmark leaderboard are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11928
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Zhou, Xiaolin
Yuan, Aojie
Luo, Zheng
Ling, Zipeng
Pan, Xixiao
Gao, Yicheng
Zhang, Haiyue
Li, Jiate
Jiang, Shuli
Wang, Prince Zizhuang
Zhu, Zixuan
Liu, Jinbo
Rossi, Ryan A.
Wei, Hua
Hu, Xiyang
Artificial Intelligence
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user typos propagate into hallucinated tool names, a misconfigured request timeout can stall an agent indefinitely, and duplicate tool names across servers can freeze an SDK. We study these failures as a sim-to-real gap in the tool-use partially observable Markov decision process (POMDP), where deployment noise enters through the observation, action space, reward-relevant metadata, or transition dynamics. We introduce RobustBench-TC, a benchmark with 22 perturbation types organized by these four POMDP components, each grounded in a verified GitHub issue or documented tool-calling failure. Across 21 models from 1.5B to 32B parameters (including the closed-source o4-mini), the robustness profile is sharply uneven: observation perturbations reduce accuracy by less than 5%, while reward-relevant and transition perturbations reduce accuracy by roughly 40% and 30%, respectively; scale alone does not close these gaps. We then propose ToolRL-DR, a domain-randomization reinforcement learning (RL) recipe that trains a tool-use agent on perturbation-augmented trajectories spanning the three statically encodable POMDP components. On a 3B backbone, ToolRL-DR-Full retains roughly three-quarters of clean accuracy and reaches an aggregate perturbed accuracy comparable to open-source 14B function-calling baselines while substantially narrowing the gap to o4-mini. It closes approximately 27% of the Transition gap despite never seeing transition perturbations in training, suggesting that RL on adversarial static tool-use inputs induces a more persistent retry policy that transfers to unseen runtime failures. The dataset, code and benchmark leaderboard are publicly available.
title When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2605.11928