TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Zhaoyang, Hu, Jiarui, Jiang, Xingyu, Zou, Pengyu, Li, Han, Peng, Chao, O'Hearn, Peter, Barr, Earl T., Harman, Mark, Sarro, Federica, Ye, He
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918516526940160
author Chu, Zhaoyang
Hu, Jiarui
Jiang, Xingyu
Zou, Pengyu
Li, Han
Peng, Chao
O'Hearn, Peter
Barr, Earl T.
Harman, Mark
Sarro, Federica
Ye, He
author_facet Chu, Zhaoyang
Hu, Jiarui
Jiang, Xingyu
Zou, Pengyu
Li, Han
Peng, Chao
O'Hearn, Peter
Barr, Earl T.
Harman, Mark
Sarro, Federica
Ye, He
contents We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 18 real-world categories, ranging from short everyday operations to workflows exceeding 50 steps, and covering 1,280 unique commands. From these, we curate a Verified subset of 200 representative, manually reviewed tasks. Comprehensive benchmarking on TerminalWorld-Verified across eight frontier models and six agents reveals that current systems still struggle with authentic terminal workflows, achieving a maximum pass rate of only 62.5%. Moreover, TerminalWorld captures real-world terminal capabilities distinct from existing expert-curated benchmarks (e.g., Terminal-Bench), with only a weak correlation to their scores (Pearson r=0.20). The automated engine makes TerminalWorld authentic and scalable by construction, enabling it to evaluate agents in real-world terminal environments as developer practices evolve. Data and code are available at https://github.com/EuniAI/TerminalWorld.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22535
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Chu, Zhaoyang
Hu, Jiarui
Jiang, Xingyu
Zou, Pengyu
Li, Han
Peng, Chao
O'Hearn, Peter
Barr, Earl T.
Harman, Mark
Sarro, Federica
Ye, He
Artificial Intelligence
We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 18 real-world categories, ranging from short everyday operations to workflows exceeding 50 steps, and covering 1,280 unique commands. From these, we curate a Verified subset of 200 representative, manually reviewed tasks. Comprehensive benchmarking on TerminalWorld-Verified across eight frontier models and six agents reveals that current systems still struggle with authentic terminal workflows, achieving a maximum pass rate of only 62.5%. Moreover, TerminalWorld captures real-world terminal capabilities distinct from existing expert-curated benchmarks (e.g., Terminal-Bench), with only a weak correlation to their scores (Pearson r=0.20). The automated engine makes TerminalWorld authentic and scalable by construction, enabling it to evaluate agents in real-world terminal environments as developer practices evolve. Data and code are available at https://github.com/EuniAI/TerminalWorld.
title TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2605.22535