CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Yusong, Wang, Haiyang, Wu, Shuzhe, Fan, Lue, Pan, Feiyang, Zhao, Sanyuan, Tu, Dandan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911441093656576
author Lin, Yusong
Wang, Haiyang
Wu, Shuzhe
Fan, Lue
Pan, Feiyang
Zhao, Sanyuan
Tu, Dandan
author_facet Lin, Yusong
Wang, Haiyang
Wu, Shuzhe
Fan, Lue
Pan, Feiyang
Zhao, Sanyuan
Tu, Dandan
contents Agentic coding requires agents to effectively interact with runtime environments, e.g., command line interfaces (CLI), so as to complete tasks like resolving dependency issues, fixing system problems, etc. But it remains underexplored how such environment-intensive tasks can be obtained at scale to enhance agents' capabilities. To address this, based on an analogy between the Dockerfile and the agentic task, we propose to employ agents to simulate and explore environment histories, guided by execution feedback. By tracing histories of a healthy environment, its state can be inverted to an earlier one with runtime failures, from which a task can be derived by packing the buggy state and the corresponding error messages. With our method, named CLI-Gym, a total of 1,655 environment-intensive tasks are derived, being the largest collection of its kind. Moreover, with curated successful trajectories, our fine-tuned model, named LiberCoder, achieves substantial absolute improvements of +21.1% (to 46.1%) on Terminal-Bench, outperforming various strong baselines. To our knowledge, this is the first public pipeline for scalable derivation of environment-intensive tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10999
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
Lin, Yusong
Wang, Haiyang
Wu, Shuzhe
Fan, Lue
Pan, Feiyang
Zhao, Sanyuan
Tu, Dandan
Artificial Intelligence
Agentic coding requires agents to effectively interact with runtime environments, e.g., command line interfaces (CLI), so as to complete tasks like resolving dependency issues, fixing system problems, etc. But it remains underexplored how such environment-intensive tasks can be obtained at scale to enhance agents' capabilities. To address this, based on an analogy between the Dockerfile and the agentic task, we propose to employ agents to simulate and explore environment histories, guided by execution feedback. By tracing histories of a healthy environment, its state can be inverted to an earlier one with runtime failures, from which a task can be derived by packing the buggy state and the corresponding error messages. With our method, named CLI-Gym, a total of 1,655 environment-intensive tasks are derived, being the largest collection of its kind. Moreover, with curated successful trajectories, our fine-tuned model, named LiberCoder, achieves substantial absolute improvements of +21.1% (to 46.1%) on Terminal-Bench, outperforming various strong baselines. To our knowledge, this is the first public pipeline for scalable derivation of environment-intensive tasks.
title CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
topic Artificial Intelligence
url https://arxiv.org/abs/2602.10999