HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
Fuente:
arXiv
Saved in:
| Main Authors: | Anokhin, Petr, Khalikov, Roman, Rebrikov, Stefan, Volkov, Viktor, Sorokin, Artyom, Bissonnette, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
by: Kuratov, Yuri, et al.
Published: (2024)
by: Kuratov, Yuri, et al.
Published: (2024)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
by: Kuratov, Yuri, et al.
Published: (2024)
by: Kuratov, Yuri, et al.
Published: (2024)
AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
by: Anokhin, Petr, et al.
Published: (2024)
by: Anokhin, Petr, et al.
Published: (2024)
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
by: Monti, Sebastiano, et al.
Published: (2026)
by: Monti, Sebastiano, et al.
Published: (2026)
MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
by: Hill, William, et al.
Published: (2023)
by: Hill, William, et al.
Published: (2023)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making
by: Grady, Thomas, et al.
Published: (2026)
by: Grady, Thomas, et al.
Published: (2026)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
by: Imajuku, Yuki, et al.
Published: (2025)
by: Imajuku, Yuki, et al.
Published: (2025)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
by: Zhang, Yinger, et al.
Published: (2026)
by: Zhang, Yinger, et al.
Published: (2026)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
by: He, Muyu, et al.
Published: (2026)
by: He, Muyu, et al.
Published: (2026)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
by: Yu, Chengjun, et al.
Published: (2026)
by: Yu, Chengjun, et al.
Published: (2026)
Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning
by: Chen, Dillon Z., et al.
Published: (2026)
by: Chen, Dillon Z., et al.
Published: (2026)
Intrinsic Stability Limits of Autoregressive Reasoning: Structural Consequences for Long-Horizon Execution
by: Liao, Hsien-Jyh
Published: (2026)
by: Liao, Hsien-Jyh
Published: (2026)
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
by: Xie, Sixiong, et al.
Published: (2026)
by: Xie, Sixiong, et al.
Published: (2026)
Verifiable Benchmarking of Long-Horizon Spatial Biology
by: Diks, Ian, et al.
Published: (2026)
by: Diks, Ian, et al.
Published: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
by: Song, Zhiheng, et al.
Published: (2026)
by: Song, Zhiheng, et al.
Published: (2026)
Beyond Entangled Planning: Task-Decoupled Planning for Long-Horizon Agents
by: Li, Yunfan, et al.
Published: (2026)
by: Li, Yunfan, et al.
Published: (2026)
MGRegBench: A Novel Benchmark Dataset with Anatomical Landmarks for Mammography Image Registration
by: Krasnova, Svetlana, et al.
Published: (2025)
by: Krasnova, Svetlana, et al.
Published: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
by: Wang, Qinsi, et al.
Published: (2026)
by: Wang, Qinsi, et al.
Published: (2026)
STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks
by: Lobo, ELita, et al.
Published: (2026)
by: Lobo, ELita, et al.
Published: (2026)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
Beyond One World: Benchmarking Super Heros in Role-Playing Across Multiversal Contexts
by: Ngokpol, Perapard, et al.
Published: (2025)
by: Ngokpol, Perapard, et al.
Published: (2025)
CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
by: Gao, Huan-ang, et al.
Published: (2025)
by: Gao, Huan-ang, et al.
Published: (2025)
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
by: Pushkin, Denys, et al.
Published: (2026)
by: Pushkin, Denys, et al.
Published: (2026)
From Real World to Logic and Back: Learning Generalizable Relational Concepts For Long Horizon Robot Planning
by: Shah, Naman, et al.
Published: (2024)
by: Shah, Naman, et al.
Published: (2024)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
by: Orlanski, Gabriel, et al.
Published: (2026)
by: Orlanski, Gabriel, et al.
Published: (2026)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
by: Cai, Yanan, et al.
Published: (2025)
by: Cai, Yanan, et al.
Published: (2025)
Spatially Grounded Long-Horizon Task Planning in the Wild
by: Jung, Sehun, et al.
Published: (2026)
by: Jung, Sehun, et al.
Published: (2026)
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
by: Chandwani, Abhishek, et al.
Published: (2026)
by: Chandwani, Abhishek, et al.
Published: (2026)
MolRecBench-Wild: A Real-World Benchmark for Optical Chemical Structure Recognition
by: Yang, Haote, et al.
Published: (2026)
by: Yang, Haote, et al.
Published: (2026)
DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant
by: Sorokin, Lev, et al.
Published: (2026)
by: Sorokin, Lev, et al.
Published: (2026)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Similar Items
-
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
by: Kuratov, Yuri, et al.
Published: (2024) -
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
by: Kuratov, Yuri, et al.
Published: (2024) -
AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
by: Anokhin, Petr, et al.
Published: (2024) -
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
by: Monti, Sebastiano, et al.
Published: (2026) -
MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
by: Hill, William, et al.
Published: (2023)