Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Zihao, Wang, Hongru, Liu, Zeming, Wang, Xinyi, Zhu, Xiangrong, Guo, Yuhang, Lin, Wei, Pan, Jeff Z., Wang, Yunhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917515443044352
author Cheng, Zihao
Wang, Hongru
Liu, Zeming
Wang, Xinyi
Zhu, Xiangrong
Guo, Yuhang
Lin, Wei
Pan, Jeff Z.
Wang, Yunhong
author_facet Cheng, Zihao
Wang, Hongru
Liu, Zeming
Wang, Xinyi
Zhu, Xiangrong
Guo, Yuhang
Lin, Wei
Pan, Jeff Z.
Wang, Yunhong
contents Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-quality training data. Existing approaches bootstrap from partial sources such as human-defined seeds or GitHub repositories to instantiate one component and then complete the rest, producing tasks confined to narrow seed distributions, environments misaligned with task semantics, and inefficient trajectories from unguided exploration. To address these limitations, we introduce Terminal-World, a fully automated pipeline that uses agent skills as the central synthesis primitive, which jointly encode what to accomplish, when to apply (preconditions and environment state), and how to execute, enabling task instructions, environments, and teacher trajectories to be co-derived. To further broaden the synthesis space, Terminal-World composes skills into skill teams and skill graphs for multi-role and cross-domain task synthesis. Using this pipeline, we construct 5,723 training environments and train Terminal-World-8B/14B/32B, evaluated across 6 benchmarks where the Terminal-World series consistently outperforms terminal-agent baselines. Notably, using the same teacher model and only 1.2% of the training data, Terminal-World-32B surpasses Nemotron-Terminal-32B on Terminal-Bench 2.0 by +4.5 Pass@1 (31.5) and achieves 43.8 Pass@3.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20876
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
Cheng, Zihao
Wang, Hongru
Liu, Zeming
Wang, Xinyi
Zhu, Xiangrong
Guo, Yuhang
Lin, Wei
Pan, Jeff Z.
Wang, Yunhong
Computation and Language
Artificial Intelligence
Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-quality training data. Existing approaches bootstrap from partial sources such as human-defined seeds or GitHub repositories to instantiate one component and then complete the rest, producing tasks confined to narrow seed distributions, environments misaligned with task semantics, and inefficient trajectories from unguided exploration. To address these limitations, we introduce Terminal-World, a fully automated pipeline that uses agent skills as the central synthesis primitive, which jointly encode what to accomplish, when to apply (preconditions and environment state), and how to execute, enabling task instructions, environments, and teacher trajectories to be co-derived. To further broaden the synthesis space, Terminal-World composes skills into skill teams and skill graphs for multi-role and cross-domain task synthesis. Using this pipeline, we construct 5,723 training environments and train Terminal-World-8B/14B/32B, evaluated across 6 benchmarks where the Terminal-World series consistently outperforms terminal-agent baselines. Notably, using the same teacher model and only 1.2% of the training data, Terminal-World-32B surpasses Nemotron-Terminal-32B on Terminal-Bench 2.0 by +4.5 Pass@1 (31.5) and achieves 43.8 Pass@3.
title Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.20876