TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Zhaoyang, Hu, Jiarui, Jiang, Xingyu, Zou, Pengyu, Li, Han, Peng, Chao, O'Hearn, Peter, Barr, Earl T., Harman, Mark, Sarro, Federica, Ye, He |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
by: Harman, Mark, et al.
Published: (2025)
by: Harman, Mark, et al.
Published: (2025)
Non-Termination Proving: 100 Million LoC and Beyond
by: Vanegue, Julien, et al.
Published: (2025)
by: Vanegue, Julien, et al.
Published: (2025)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
HotBugs.jar: A Benchmark of Hot Fixes for Time-Critical Bugs
by: Hanna, Carol, et al.
Published: (2025)
by: Hanna, Carol, et al.
Published: (2025)
Generative AI for Testing of Autonomous Driving Systems: A Survey
by: Song, Qunying, et al.
Published: (2025)
by: Song, Qunying, et al.
Published: (2025)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
by: Sultan, Oren, et al.
Published: (2026)
by: Sultan, Oren, et al.
Published: (2026)
Fairness Improvement with Multiple Protected Attributes: How Far Are We?
by: Chen, Zhenpeng, et al.
Published: (2023)
by: Chen, Zhenpeng, et al.
Published: (2023)
EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents
by: Guo, Yaoqi, et al.
Published: (2026)
by: Guo, Yaoqi, et al.
Published: (2026)
NEWSAGENT: Benchmarking Multimodal Agents as Journalists with Real-World Newswriting Tasks
by: Chien, Yen-Che, et al.
Published: (2025)
by: Chien, Yen-Che, et al.
Published: (2025)
Observation of manganese crusts recovered by the USNS Bartlett in 1973 over the East Pacific Rise
by: Melson, William G, et al.
Published: (1986)
by: Melson, William G, et al.
Published: (1986)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
by: Merrill, Mike A., et al.
Published: (2026)
by: Merrill, Mike A., et al.
Published: (2026)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
Fairness Testing: A Comprehensive Survey and Analysis of Trends
by: Chen, Zhenpeng, et al.
Published: (2022)
by: Chen, Zhenpeng, et al.
Published: (2022)
No More, No Less: Task Alignment in Terminal Agents
by: Mavali, Sina, et al.
Published: (2026)
by: Mavali, Sina, et al.
Published: (2026)
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
by: Peng, Xiaoxuan, et al.
Published: (2026)
by: Peng, Xiaoxuan, et al.
Published: (2026)
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
by: Chen, Ying, et al.
Published: (2026)
by: Chen, Ying, et al.
Published: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
Geochemistry and minerals at DSDP Leg 65 Holes
by: Flower, Martin F J, et al.
Published: (1983)
by: Flower, Martin F J, et al.
Published: (1983)
(Table 1) Geochemistry and minerals of glass selvedges at DSDP Leg 65 Holes
by: Flower, Martin F J, et al.
Published: (1983)
by: Flower, Martin F J, et al.
Published: (1983)
(Table 2) Geochemistry and minerals of whole-rock samples analyzed as fused beads at DSDP Leg 65 Holes
by: Flower, Martin F J, et al.
Published: (1983)
by: Flower, Martin F J, et al.
Published: (1983)
Practitioners' Expectations on Log Anomaly Detection
by: Ma, Xiaoxue, et al.
Published: (2024)
by: Ma, Xiaoxue, et al.
Published: (2024)
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
by: Heo, Chiyeong, et al.
Published: (2026)
by: Heo, Chiyeong, et al.
Published: (2026)
Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance
by: Pinna, Giovanni, et al.
Published: (2026)
by: Pinna, Giovanni, et al.
Published: (2026)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
by: Xu, Frank F., et al.
Published: (2024)
by: Xu, Frank F., et al.
Published: (2024)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
by: Song, Zhiheng, et al.
Published: (2026)
by: Song, Zhiheng, et al.
Published: (2026)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)
by: Meng, Jinxiang, et al.
Published: (2026)
(Table 2) Stratigraphic locations for glassy objects in DSDP Hole 5-32
by: Melson, William G, et al.
Published: (1988)
by: Melson, William G, et al.
Published: (1988)
(Table 3) Basement analyses from DSDP Hole 5-32
by: Melson, William G, et al.
Published: (1988)
by: Melson, William G, et al.
Published: (1988)
Logic.py: Bridging the Gap between LLMs and Constraint Solvers
by: Kesseli, Pascal, et al.
Published: (2025)
by: Kesseli, Pascal, et al.
Published: (2025)
(Table 3) Spherules analyses from DSDP Hole 5-32
by: Melson, William G, et al.
Published: (1988)
by: Melson, William G, et al.
Published: (1988)
Spherules and basement analyses from DSDP Hole 5-32
by: Melson, William G, et al.
Published: (1988)
by: Melson, William G, et al.
Published: (1988)
Termination of Real Linear Loops
by: Neumann, Eike, et al.
Published: (2026)
by: Neumann, Eike, et al.
Published: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
Studies on the Termination Behavior of Styrenic Free‐Radicals
by: Wei Zhijie, et al.
Published: (2024)
by: Wei Zhijie, et al.
Published: (2024)
EricStorm. 2024. Nationalism: A World History. Princeton University Press. 491pp. £35.00 (hbk).
by: Jonathan Hearn
Published: (2026)
by: Jonathan Hearn
Published: (2026)
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
by: Chi, Yizhe, et al.
Published: (2026)
by: Chi, Yizhe, et al.
Published: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)
by: Lee, Hwiwon, et al.
Published: (2025)
Similar Items
-
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
by: Harman, Mark, et al.
Published: (2025) -
Non-Termination Proving: 100 Million LoC and Beyond
by: Vanegue, Julien, et al.
Published: (2025) -
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026) -
HotBugs.jar: A Benchmark of Hot Fixes for Time-Critical Bugs
by: Hanna, Carol, et al.
Published: (2025) -
Generative AI for Testing of Autonomous Driving Systems: A Survey
by: Song, Qunying, et al.
Published: (2025)