UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Luo, Haotian, Zhang, Huaisong, Zhang, Xuelin, Wang, Haoyu, Qin, Zeyu, Lu, Wenjie, Ma, Guozheng, He, Haiying, Xie, Yingsha, Zhou, Qiyang, Hu, Zixuan, Mi, Hongze, Wang, Yibo, Tan, Naiqiang, Chen, Hong, Fung, Yi R., Yuan, Chun, Shen, Li
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908559335227392
author Luo, Haotian
Zhang, Huaisong
Zhang, Xuelin
Wang, Haoyu
Qin, Zeyu
Lu, Wenjie
Ma, Guozheng
He, Haiying
Xie, Yingsha
Zhou, Qiyang
Hu, Zixuan
Mi, Hongze
Wang, Yibo
Tan, Naiqiang
Chen, Hong
Fung, Yi R.
Yuan, Chun
Shen, Li
author_facet Luo, Haotian
Zhang, Huaisong
Zhang, Xuelin
Wang, Haoyu
Qin, Zeyu
Lu, Wenjie
Ma, Guozheng
He, Haiying
Xie, Yingsha
Zhou, Qiyang
Hu, Zixuan
Mi, Hongze
Wang, Yibo
Tan, Naiqiang
Chen, Hong
Fung, Yi R.
Yuan, Chun
Shen, Li
contents Autonomous agents have recently achieved remarkable progress across diverse domains, yet most evaluations focus on short-horizon, fully observable tasks. In contrast, many critical real-world tasks, such as large-scale software development, commercial investment, and scientific discovery, unfold in long-horizon and partially observable scenarios where success hinges on sustained reasoning, planning, memory management, and tool use. Existing benchmarks rarely capture these long-horizon challenges, leaving a gap in systematic evaluation. To bridge this gap, we introduce \textbf{UltraHorizon} a novel benchmark that measures the foundational capabilities essential for complex real-world challenges. We use exploration as a unifying task across three distinct environments to validate these core competencies. Agents are designed in long-horizon discovery tasks where they must iteratively uncover hidden rules through sustained reasoning, planning, memory and tools management, and interaction with environments. Under the heaviest scale setting, trajectories average \textbf{200k+} tokens and \textbf{400+} tool calls, whereas in standard configurations they still exceed \textbf{35k} tokens and involve more than \textbf{60} tool calls on average. Our extensive experiments reveal that LLM-agents consistently underperform in these settings, whereas human participants achieve higher scores, underscoring a persistent gap in agents' long-horizon abilities. We also observe that simple scaling fails in our task. To better illustrate the failure of agents, we conduct an in-depth analysis of collected trajectories. We identify eight types of errors and attribute them to two primary causes: in-context locking and functional fundamental capability gaps. \href{https://github.com/StarDewXXX/UltraHorizon}{Our code will be available here.}
format Preprint
id arxiv_https___arxiv_org_abs_2509_21766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
Luo, Haotian
Zhang, Huaisong
Zhang, Xuelin
Wang, Haoyu
Qin, Zeyu
Lu, Wenjie
Ma, Guozheng
He, Haiying
Xie, Yingsha
Zhou, Qiyang
Hu, Zixuan
Mi, Hongze
Wang, Yibo
Tan, Naiqiang
Chen, Hong
Fung, Yi R.
Yuan, Chun
Shen, Li
Artificial Intelligence
Computation and Language
Autonomous agents have recently achieved remarkable progress across diverse domains, yet most evaluations focus on short-horizon, fully observable tasks. In contrast, many critical real-world tasks, such as large-scale software development, commercial investment, and scientific discovery, unfold in long-horizon and partially observable scenarios where success hinges on sustained reasoning, planning, memory management, and tool use. Existing benchmarks rarely capture these long-horizon challenges, leaving a gap in systematic evaluation. To bridge this gap, we introduce \textbf{UltraHorizon} a novel benchmark that measures the foundational capabilities essential for complex real-world challenges. We use exploration as a unifying task across three distinct environments to validate these core competencies. Agents are designed in long-horizon discovery tasks where they must iteratively uncover hidden rules through sustained reasoning, planning, memory and tools management, and interaction with environments. Under the heaviest scale setting, trajectories average \textbf{200k+} tokens and \textbf{400+} tool calls, whereas in standard configurations they still exceed \textbf{35k} tokens and involve more than \textbf{60} tool calls on average. Our extensive experiments reveal that LLM-agents consistently underperform in these settings, whereas human participants achieve higher scores, underscoring a persistent gap in agents' long-horizon abilities. We also observe that simple scaling fails in our task. To better illustrate the failure of agents, we conduct an in-depth analysis of collected trajectories. We identify eight types of errors and attribute them to two primary causes: in-context locking and functional fundamental capability gaps. \href{https://github.com/StarDewXXX/UltraHorizon}{Our code will be available here.}
title UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.21766