Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.14504 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913135831547904 |
|---|---|
| author | Zhu, Zilin Guo, Longteng Mei, Yanghong Pang, Bowen Zhang, Zongxun He, Xingjian Ji, Ruyi Liu, Jing |
| author_facet | Zhu, Zilin Guo, Longteng Mei, Yanghong Pang, Bowen Zhang, Zongxun He, Xingjian Ji, Ruyi Liu, Jing |
| contents | Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed task categories. We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning. We further propose HoloMind, a VLM-driven agent with a DAG-based long-horizon hierarchical planner, a Multimodal Spatial Memory for persistent world modeling, an Episodic Memory for experience reuse, and a global Critic for reflective supervision. Experiments with GPT-5 and Qwen3-VL models show that HoloMind substantially improves long-horizon performance while reducing reliance on model scale. Even top models achieve only 59% goal completion and 16% full-task success, underscoring the difficulty of LongAct and the need for stronger long-horizon planning in embodied agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_14504 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution Zhu, Zilin Guo, Longteng Mei, Yanghong Pang, Bowen Zhang, Zongxun He, Xingjian Ji, Ruyi Liu, Jing Artificial Intelligence Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed task categories. We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning. We further propose HoloMind, a VLM-driven agent with a DAG-based long-horizon hierarchical planner, a Multimodal Spatial Memory for persistent world modeling, an Episodic Memory for experience reuse, and a global Critic for reflective supervision. Experiments with GPT-5 and Qwen3-VL models show that HoloMind substantially improves long-horizon performance while reducing reliance on model scale. Even top models achieve only 59% goal completion and 16% full-task success, underscoring the difficulty of LongAct and the need for stronger long-horizon planning in embodied agents. |
| title | When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2605.14504 |