Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Siyuan, Li, Shiyang, Liu, Xin, Liu, Tianyi, Li, Yixiao, Shi, Zhan, Zhang, Zixuan, Wang, Zilong, Yin, Qingyu, Chen, Jianshu, Zhao, Tuo, Yin, Bing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908954195394560
author Xu, Siyuan
Li, Shiyang
Liu, Xin
Liu, Tianyi
Li, Yixiao
Shi, Zhan
Zhang, Zixuan
Wang, Zilong
Yin, Qingyu
Chen, Jianshu
Zhao, Tuo
Yin, Bing
author_facet Xu, Siyuan
Li, Shiyang
Liu, Xin
Liu, Tianyi
Li, Yixiao
Shi, Zhan
Zhang, Zixuan
Wang, Zilong
Yin, Qingyu
Chen, Jianshu
Zhao, Tuo
Yin, Bing
contents Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable environments that support reward-checkable online rollouts. We propose COVERT, a two-stage pipeline that first generates reliable base tool-use trajectories through self-evolving synthesis with multi-level validation, and then applies oracle-preserving augmentations that systematically increase environmental complexity. These augmentations introduce distractor tools, indirect or ambiguous user queries, and noisy, multi-format, or erroneous tool outputs, while strictly preserving oracle tool calls and final answers as ground truth. This design enables automatic reward computation via reference matching for standard cases and lightweight judge-assisted verification for special behaviors such as error detection, supporting RL optimization of tool-calling policies. On Qwen2.5-Instruct-14B, COVERT-RL improves overall accuracy on BFCL v3 from 56.5 to 59.9 and on ACEBench from 53.0 to 59.3, with minimal regressions on general-ability benchmarks; when stacked on SFT, it further reaches 62.1 and 61.8, confirming additive gains. These results suggest that oracle-preserving synthetic environments offer a practical RL refinement stage, complementary to SFT, for improving tool-use robustness under ambiguity and unreliable tool feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
Xu, Siyuan
Li, Shiyang
Liu, Xin
Liu, Tianyi
Li, Yixiao
Shi, Zhan
Zhang, Zixuan
Wang, Zilong
Yin, Qingyu
Chen, Jianshu
Zhao, Tuo
Yin, Bing
Artificial Intelligence
Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable environments that support reward-checkable online rollouts. We propose COVERT, a two-stage pipeline that first generates reliable base tool-use trajectories through self-evolving synthesis with multi-level validation, and then applies oracle-preserving augmentations that systematically increase environmental complexity. These augmentations introduce distractor tools, indirect or ambiguous user queries, and noisy, multi-format, or erroneous tool outputs, while strictly preserving oracle tool calls and final answers as ground truth. This design enables automatic reward computation via reference matching for standard cases and lightweight judge-assisted verification for special behaviors such as error detection, supporting RL optimization of tool-calling policies. On Qwen2.5-Instruct-14B, COVERT-RL improves overall accuracy on BFCL v3 from 56.5 to 59.9 and on ACEBench from 53.0 to 59.3, with minimal regressions on general-ability benchmarks; when stacked on SFT, it further reaches 62.1 and 61.8, confirming additive gains. These results suggest that oracle-preserving synthetic environments offer a practical RL refinement stage, complementary to SFT, for improving tool-use robustness under ambiguity and unreliable tool feedback.
title Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2604.09813