Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Weihua, Gong, Hailei, Ling, Zhan, Liu, Kang, Shen, Lingfeng, Yao, Xuesong, Xu, Yufei, Shi, Dingyuan, Yang, Yiming, Chen, Jiecao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
by: Lu, Miao, et al.
Published: (2025)
by: Lu, Miao, et al.
Published: (2025)
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025)
by: Ling, Zhan, et al.
Published: (2025)
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
by: Lai, Hanyu, et al.
Published: (2025)
by: Lai, Hanyu, et al.
Published: (2025)
ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Booster Gym: An End-to-End Reinforcement Learning Framework for Humanoid Robot Locomotion
by: Wang, Yushi, et al.
Published: (2025)
by: Wang, Yushi, et al.
Published: (2025)
Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving
by: Ge, Junhao, et al.
Published: (2025)
by: Ge, Junhao, et al.
Published: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025)
by: Chen, Xuesong, et al.
Published: (2025)
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
by: Yao, Wenhao, et al.
Published: (2026)
by: Yao, Wenhao, et al.
Published: (2026)
End-to-End Learning for Partially-Observed Time Series with PyPOTS
by: Du, Wenjie, et al.
Published: (2026)
by: Du, Wenjie, et al.
Published: (2026)
SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
by: Lee, Sangmin, et al.
Published: (2025)
by: Lee, Sangmin, et al.
Published: (2025)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving
by: Sun, Jiangxin, et al.
Published: (2026)
by: Sun, Jiangxin, et al.
Published: (2026)
End to End Collaborative Synthetic Data Generation
by: Pentyala, Sikha, et al.
Published: (2024)
by: Pentyala, Sikha, et al.
Published: (2024)
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
by: Fan, Zhiyuan, et al.
Published: (2026)
by: Fan, Zhiyuan, et al.
Published: (2026)
X-MOBILITY: End-To-End Generalizable Navigation via World Modeling
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Bits-to-Photon: End-to-End Learned Scalable Point Cloud Compression for Direct Rendering
by: Hu, Yueyu, et al.
Published: (2024)
by: Hu, Yueyu, et al.
Published: (2024)
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python?
by: Gong, Jianian, et al.
Published: (2024)
by: Gong, Jianian, et al.
Published: (2024)
SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
by: Zheng, Peiru, et al.
Published: (2025)
by: Zheng, Peiru, et al.
Published: (2025)
Minimizing End-to-End Latency for Joint Source-Channel Coding Systems
by: Chi, Kaiyi, et al.
Published: (2024)
by: Chi, Kaiyi, et al.
Published: (2024)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
FalconApp: Rapid iPhone Deployment of End-to-End Perception via Automatically Labeled Synthetic Data
by: Miao, Yan, et al.
Published: (2026)
by: Miao, Yan, et al.
Published: (2026)
Embodied Cognition Augmented End2End Autonomous Driving
by: Niu, Ling, et al.
Published: (2025)
by: Niu, Ling, et al.
Published: (2025)
MAP: End-to-End Autonomous Driving with Map-Assisted Planning
by: Yin, Huilin, et al.
Published: (2025)
by: Yin, Huilin, et al.
Published: (2025)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
Sim-to-Real gap in RL: Use Case with TIAGo and Isaac Sim/Gym
by: Albardaner, Jaume, et al.
Published: (2024)
by: Albardaner, Jaume, et al.
Published: (2024)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
by: Li, Weizhen, et al.
Published: (2025)
by: Li, Weizhen, et al.
Published: (2025)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
by: Lu, Dekun, et al.
Published: (2025)
by: Lu, Dekun, et al.
Published: (2025)
Graph Transformer Networks for Accurate Band Structure Prediction: An End-to-End Approach
by: Gong, Weiyi, et al.
Published: (2024)
by: Gong, Weiyi, et al.
Published: (2024)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
by: Zheng, Peiru, et al.
Published: (2024)
by: Zheng, Peiru, et al.
Published: (2024)
FormalASR: End-to-End Spoken Chinese to Formal Text
by: Ning, Wanyi, et al.
Published: (2026)
by: Ning, Wanyi, et al.
Published: (2026)
Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
by: Khurram, Aleesha, et al.
Published: (2025)
by: Khurram, Aleesha, et al.
Published: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
Similar Items
-
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
by: Lu, Miao, et al.
Published: (2025) -
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025) -
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025) -
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
by: Ye, Junjie, et al.
Published: (2025) -
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025)