When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xiaolin, Yuan, Aojie, Luo, Zheng, Ling, Zipeng, Pan, Xixiao, Gao, Yicheng, Zhang, Haiyue, Li, Jiate, Jiang, Shuli, Wang, Prince Zizhuang, Zhu, Zixuan, Liu, Jinbo, Rossi, Ryan A., Wei, Hua, Hu, Xiyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
Auditable Agents
by: Nian, Yi, et al.
Published: (2026)
by: Nian, Yi, et al.
Published: (2026)
Counterfactual Trace Auditing of LLM Agent Skills
by: Zhou, Xiaolin, et al.
Published: (2026)
by: Zhou, Xiaolin, et al.
Published: (2026)
Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
OpenTable-R1: A Reinforcement Learning Augmented Tool Agent for Open-Domain Table Question Answering
by: Qiu, Zipeng
Published: (2025)
by: Qiu, Zipeng
Published: (2025)
Sim-to-Real gap in RL: Use Case with TIAGo and Isaac Sim/Gym
by: Albardaner, Jaume, et al.
Published: (2024)
by: Albardaner, Jaume, et al.
Published: (2024)
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
HMAR: Hierarchical Modality-Aware Expert and Dynamic Routing Medical Image Retrieval Architecture
by: Yuan, Aojie
Published: (2026)
by: Yuan, Aojie
Published: (2026)
AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
Gen-DFL: Decision-Focused Generative Learning for Robust Decision Making
by: Wang, Prince Zizhuang, et al.
Published: (2025)
by: Wang, Prince Zizhuang, et al.
Published: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
by: Gong, Guangyu, et al.
Published: (2026)
by: Gong, Guangyu, et al.
Published: (2026)
"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
by: Li, Jiate, et al.
Published: (2026)
by: Li, Jiate, et al.
Published: (2026)
Vero: An Open RL Recipe for General Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2026)
by: Sarch, Gabriel, et al.
Published: (2026)
AGNNCert: Defending Graph Neural Networks against Arbitrary Perturbations with Deterministic Certification
by: Li, Jiate, et al.
Published: (2025)
by: Li, Jiate, et al.
Published: (2025)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
by: Zhou, Xiaolin, et al.
Published: (2026)
by: Zhou, Xiaolin, et al.
Published: (2026)
Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
Spectroscopic Quantification of Plasma Free Hemoglobin Based on Paired Domain Adaptation and Orthogonality Constraints
by: Haiyue Lv, et al.
Published: (2026)
by: Haiyue Lv, et al.
Published: (2026)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
by: Maddukuri, Abhiram, et al.
Published: (2025)
by: Maddukuri, Abhiram, et al.
Published: (2025)
When Deepfake Detection Meets Graph Neural Network:a Unified and Lightweight Learning Framework
by: Liu, Haoyu, et al.
Published: (2025)
by: Liu, Haoyu, et al.
Published: (2025)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
by: Chang, Qikai, et al.
Published: (2025)
by: Chang, Qikai, et al.
Published: (2025)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
by: Dai, Fangqi, et al.
Published: (2025)
by: Dai, Fangqi, et al.
Published: (2025)
SimFuzz: Similarity-guided Block-level Mutation for RISC-V Processor Fuzzing
by: Lyu, Hao, et al.
Published: (2026)
by: Lyu, Hao, et al.
Published: (2026)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
Federated Learning with Incomplete Data: When to Use Complete Cases and When to Weight
by: Vazquez, Jesus E., et al.
Published: (2026)
by: Vazquez, Jesus E., et al.
Published: (2026)
A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
Enhanced cooling via self-Kerr nonlinearity in cavity-magnomechanical system
by: Xu, Jiate, et al.
Published: (2025)
by: Xu, Jiate, et al.
Published: (2025)
Practicable Black-box Evasion Attacks on Link Prediction in Dynamic Graphs -- A Graph Sequential Embedding Method
by: Li, Jiate, et al.
Published: (2024)
by: Li, Jiate, et al.
Published: (2024)
Closing the Sim2Real Performance Gap in RL
by: Anand, Akhil S, et al.
Published: (2025)
by: Anand, Akhil S, et al.
Published: (2025)
Bridging the Sim-to-Real Gap from the Information Bottleneck Perspective
by: He, Haoran, et al.
Published: (2023)
by: He, Haoran, et al.
Published: (2023)
ToRL: Scaling Tool-Integrated RL
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
RecipeGen: A Benchmark for Real-World Recipe Image Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
by: Xu, Ningning, et al.
Published: (2025)
by: Xu, Ningning, et al.
Published: (2025)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
Similar Items
-
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
by: Wang, Prince Zizhuang, et al.
Published: (2026) -
AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks
by: Wang, Prince Zizhuang, et al.
Published: (2026) -
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent
by: Wang, Prince Zizhuang, et al.
Published: (2026) -
Auditable Agents
by: Nian, Yi, et al.
Published: (2026) -
Counterfactual Trace Auditing of LLM Agent Skills
by: Zhou, Xiaolin, et al.
Published: (2026)