Saved in:
Bibliographic Details
Main Authors: Zhu, Zilin, Guo, Longteng, Mei, Yanghong, Pang, Bowen, Zhang, Zongxun, He, Xingjian, Ji, Ruyi, Liu, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.14504
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913135831547904
author Zhu, Zilin
Guo, Longteng
Mei, Yanghong
Pang, Bowen
Zhang, Zongxun
He, Xingjian
Ji, Ruyi
Liu, Jing
author_facet Zhu, Zilin
Guo, Longteng
Mei, Yanghong
Pang, Bowen
Zhang, Zongxun
He, Xingjian
Ji, Ruyi
Liu, Jing
contents Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed task categories. We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning. We further propose HoloMind, a VLM-driven agent with a DAG-based long-horizon hierarchical planner, a Multimodal Spatial Memory for persistent world modeling, an Episodic Memory for experience reuse, and a global Critic for reflective supervision. Experiments with GPT-5 and Qwen3-VL models show that HoloMind substantially improves long-horizon performance while reducing reliance on model scale. Even top models achieve only 59% goal completion and 16% full-task success, underscoring the difficulty of LongAct and the need for stronger long-horizon planning in embodied agents.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14504
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
Zhu, Zilin
Guo, Longteng
Mei, Yanghong
Pang, Bowen
Zhang, Zongxun
He, Xingjian
Ji, Ruyi
Liu, Jing
Artificial Intelligence
Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed task categories. We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning. We further propose HoloMind, a VLM-driven agent with a DAG-based long-horizon hierarchical planner, a Multimodal Spatial Memory for persistent world modeling, an Episodic Memory for experience reuse, and a global Critic for reflective supervision. Experiments with GPT-5 and Qwen3-VL models show that HoloMind substantially improves long-horizon performance while reducing reliance on model scale. Even top models achieve only 59% goal completion and 16% full-task success, underscoring the difficulty of LongAct and the need for stronger long-horizon planning in embodied agents.
title When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
topic Artificial Intelligence
url https://arxiv.org/abs/2605.14504