Generalizable End-to-End Tool-Use RL with Synthetic CodeGym

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Weihua, Gong, Hailei, Ling, Zhan, Liu, Kang, Shen, Lingfeng, Yao, Xuesong, Xu, Yufei, Shi, Dingyuan, Yang, Yiming, Chen, Jiecao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908890487062528
author Du, Weihua
Gong, Hailei
Ling, Zhan
Liu, Kang
Shen, Lingfeng
Yao, Xuesong
Xu, Yufei
Shi, Dingyuan
Yang, Yiming
Chen, Jiecao
author_facet Du, Weihua
Gong, Hailei
Ling, Zhan
Liu, Kang
Shen, Lingfeng
Yao, Xuesong
Xu, Yufei
Shi, Dingyuan
Yang, Yiming
Chen, Jiecao
contents Tool-augmented large language models (LLMs), hereafter LLM agents, leverage external tools to solve diverse tasks and interface with the real world. However, current training practices largely rely on supervised fine-tuning (SFT) over static trajectories or reinforcement learning (RL) on narrow tasks, which generalize poorly beyond development settings and lead to brittleness with new tools and unseen workflows. Because code execution reflects many structural patterns of real-world workflows, we use coding problems as a structured substrate to build tool-use agent training environments with diverse task configurations. To this end, we introduce CodeGym, a scalable framework that synthesizes diverse, verifiable, and controllable multi-turn tool-use environments for agent RL, enabling LLM agents to explore and master various workflows actively. CodeGym converts static coding problems into interactive environments by extracting atomic functions or logic into callable tools, yielding verifiable tasks that span various tool-execution workflows. Models of varying sizes and chain-of-thought configurations trained in CodeGym exhibit consistent out-of-distribution generalizability; for example, Qwen2.5-32B-Instruct achieves an absolute accuracy gain of 8.7 points on the OOD benchmark $τ$-Bench. These results highlight CodeGym as a step toward scalable general-purpose RL environments for training tool-use behaviors that align with real-world agent workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
Du, Weihua
Gong, Hailei
Ling, Zhan
Liu, Kang
Shen, Lingfeng
Yao, Xuesong
Xu, Yufei
Shi, Dingyuan
Yang, Yiming
Chen, Jiecao
Machine Learning
Artificial Intelligence
Computation and Language
Tool-augmented large language models (LLMs), hereafter LLM agents, leverage external tools to solve diverse tasks and interface with the real world. However, current training practices largely rely on supervised fine-tuning (SFT) over static trajectories or reinforcement learning (RL) on narrow tasks, which generalize poorly beyond development settings and lead to brittleness with new tools and unseen workflows. Because code execution reflects many structural patterns of real-world workflows, we use coding problems as a structured substrate to build tool-use agent training environments with diverse task configurations. To this end, we introduce CodeGym, a scalable framework that synthesizes diverse, verifiable, and controllable multi-turn tool-use environments for agent RL, enabling LLM agents to explore and master various workflows actively. CodeGym converts static coding problems into interactive environments by extracting atomic functions or logic into callable tools, yielding verifiable tasks that span various tool-execution workflows. Models of varying sizes and chain-of-thought configurations trained in CodeGym exhibit consistent out-of-distribution generalizability; for example, Qwen2.5-32B-Instruct achieves an absolute accuracy gain of 8.7 points on the OOD benchmark $τ$-Bench. These results highlight CodeGym as a step toward scalable general-purpose RL environments for training tool-use behaviors that align with real-world agent workflows.
title Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.17325