UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Haotian, Zhang, Huaisong, Zhang, Xuelin, Wang, Haoyu, Qin, Zeyu, Lu, Wenjie, Ma, Guozheng, He, Haiying, Xie, Yingsha, Zhou, Qiyang, Hu, Zixuan, Mi, Hongze, Wang, Yibo, Tan, Naiqiang, Chen, Hong, Fung, Yi R., Yuan, Chun, Shen, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
by: Mi, Hongze, et al.
Published: (2025)
by: Mi, Hongze, et al.
Published: (2025)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
by: Xie, Yingsha, et al.
Published: (2026)
by: Xie, Yingsha, et al.
Published: (2026)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution
by: Mi, Hongze, et al.
Published: (2026)
by: Mi, Hongze, et al.
Published: (2026)
Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
by: Li, Yanming, et al.
Published: (2026)
by: Li, Yanming, et al.
Published: (2026)
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026)
by: Su, Zhaochen, et al.
Published: (2026)
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
by: Alam, Samiul, et al.
Published: (2026)
by: Alam, Samiul, et al.
Published: (2026)
Beyond Short-Horizon: VQ-Memory for Robust Long-Horizon Manipulation in Non-Markovian Simulation Benchmarks
by: Wang, Honghui, et al.
Published: (2026)
by: Wang, Honghui, et al.
Published: (2026)
MARPLE: A Benchmark for Long-Horizon Inference
by: Jin, Emily, et al.
Published: (2024)
by: Jin, Emily, et al.
Published: (2024)
HiRegEx: Interactive Visual Query and Exploration of Multivariate Hierarchical Data
by: Li, Guozheng, et al.
Published: (2024)
by: Li, Guozheng, et al.
Published: (2024)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
by: Zhang, Yuanliang, et al.
Published: (2024)
by: Zhang, Yuanliang, et al.
Published: (2024)
HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue
by: Lin, Jingjie, et al.
Published: (2026)
by: Lin, Jingjie, et al.
Published: (2026)
AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
by: Jiang, Tanqiu, et al.
Published: (2026)
by: Jiang, Tanqiu, et al.
Published: (2026)
Milestone-Guided Policy Learning for Long-Horizon Language Agents
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Inconsistencies of Tsallis Cosmology within Horizon Thermodynamics and Holographic Scenarios
by: Ibarbo-Perlaza., Pedro M., et al.
Published: (2025)
by: Ibarbo-Perlaza., Pedro M., et al.
Published: (2025)
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
Why can a hydrophilic polyelectrolyte precipitate and redissolve below the critical micelle concentration of an oppositely-charged surfactant ?
by: Yong, Huaisong
Published: (2024)
by: Yong, Huaisong
Published: (2024)
UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
by: Chen, Xueyao, et al.
Published: (2026)
by: Chen, Xueyao, et al.
Published: (2026)
ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Prediction Horizons
by: Zhang, Jiawen, et al.
Published: (2023)
by: Zhang, Jiawen, et al.
Published: (2023)
Verifiable Benchmarking of Long-Horizon Spatial Biology
by: Diks, Ian, et al.
Published: (2026)
by: Diks, Ian, et al.
Published: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
by: Xu, Pengju, et al.
Published: (2025)
by: Xu, Pengju, et al.
Published: (2025)
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
by: Wang, Xinyu Jessica, et al.
Published: (2026)
by: Wang, Xinyu Jessica, et al.
Published: (2026)
What Makes Value Learning Efficient in Residual Reinforcement Learning?
by: Ma, Guozheng, et al.
Published: (2026)
by: Ma, Guozheng, et al.
Published: (2026)
Similar Items
-
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
by: Mi, Hongze, et al.
Published: (2025) -
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025) -
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
by: Xie, Yingsha, et al.
Published: (2026) -
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026) -
Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution
by: Mi, Hongze, et al.
Published: (2026)