LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Chandwani, Abhishek, Gupta, Ishan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
by: Wu, Xiyang, et al.
Published: (2026)
by: Wu, Xiyang, et al.
Published: (2026)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
by: Srininvasan, Vasundra
Published: (2026)
by: Srininvasan, Vasundra
Published: (2026)
Spatially Grounded Long-Horizon Task Planning in the Wild
by: Jung, Sehun, et al.
Published: (2026)
by: Jung, Sehun, et al.
Published: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
by: Li, Xiangyi, et al.
Published: (2026)
by: Li, Xiangyi, et al.
Published: (2026)
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
by: Han, Yi, et al.
Published: (2026)
by: Han, Yi, et al.
Published: (2026)
Learning Agent-Compatible Context Management for Long-Horizon Tasks
by: Yi, Lu, et al.
Published: (2026)
by: Yi, Lu, et al.
Published: (2026)
SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks
by: Wen, Yongyan, et al.
Published: (2024)
by: Wen, Yongyan, et al.
Published: (2024)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
by: Orlanski, Gabriel, et al.
Published: (2026)
by: Orlanski, Gabriel, et al.
Published: (2026)
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
by: Monti, Sebastiano, et al.
Published: (2026)
by: Monti, Sebastiano, et al.
Published: (2026)
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
by: Ren, Qingnan, et al.
Published: (2026)
by: Ren, Qingnan, et al.
Published: (2026)
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
by: Zhang, Linghua, et al.
Published: (2026)
by: Zhang, Linghua, et al.
Published: (2026)
Beyond Entangled Planning: Task-Decoupled Planning for Long-Horizon Agents
by: Li, Yunfan, et al.
Published: (2026)
by: Li, Yunfan, et al.
Published: (2026)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks
by: Shao, Jie-Jing, et al.
Published: (2026)
by: Shao, Jie-Jing, et al.
Published: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
ELHPlan: Efficient Long-Horizon Task Planning for Multi-Agent Collaboration
by: Ling, Shaobin, et al.
Published: (2025)
by: Ling, Shaobin, et al.
Published: (2025)
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
by: Hu, Yuyang, et al.
Published: (2026)
by: Hu, Yuyang, et al.
Published: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026)
by: Guo, Zhengkang, et al.
Published: (2026)
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
by: Kleczek, Darek, et al.
Published: (2026)
by: Kleczek, Darek, et al.
Published: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
LHManip: A Dataset for Long-Horizon Language-Grounded Manipulation Tasks in Cluttered Tabletop Environments
by: Ceola, Federico, et al.
Published: (2023)
by: Ceola, Federico, et al.
Published: (2023)
Orchestrator: Active Inference for Multi-Agent Systems in Long-Horizon Tasks
by: Beckenbauer, Lukas, et al.
Published: (2025)
by: Beckenbauer, Lukas, et al.
Published: (2025)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents
by: Chen, Zhuofu, et al.
Published: (2026)
by: Chen, Zhuofu, et al.
Published: (2026)
ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents
by: Guo, Shuhan, et al.
Published: (2026)
by: Guo, Shuhan, et al.
Published: (2026)
PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
by: Orimo, Yuki, et al.
Published: (2025)
by: Orimo, Yuki, et al.
Published: (2025)
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
by: Lei, Mingcong, et al.
Published: (2025)
by: Lei, Mingcong, et al.
Published: (2025)
CoMIC: Collaborative Memory and Insights Circulation for Long-Horizon LLM Agents in Cloud-Edge Systems
by: Wang, Yannan, et al.
Published: (2026)
by: Wang, Yannan, et al.
Published: (2026)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
by: Xu, Xinbo, et al.
Published: (2026)
by: Xu, Xinbo, et al.
Published: (2026)
LLM-Powered Knowledge Graphs for Enterprise Intelligence and Analytics
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making
by: Grady, Thomas, et al.
Published: (2026)
by: Grady, Thomas, et al.
Published: (2026)
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
by: Zhu, Zilin, et al.
Published: (2026)
by: Zhu, Zilin, et al.
Published: (2026)
Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
by: Fu, Dayuan, et al.
Published: (2025)
by: Fu, Dayuan, et al.
Published: (2025)
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
Similar Items
-
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
by: Wu, Xiyang, et al.
Published: (2026) -
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
by: Zhang, Haoran, et al.
Published: (2026) -
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
by: Srininvasan, Vasundra
Published: (2026) -
Spatially Grounded Long-Horizon Task Planning in the Wild
by: Jung, Sehun, et al.
Published: (2026) -
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)