Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuelin, Cheng, Sijie, Li, Chen, Li, Zongzhao, Huang, Yuxin, Liu, Yang, Huang, Wenbing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912972378472448
author Zhang, Yuelin
Cheng, Sijie
Li, Chen
Li, Zongzhao
Huang, Yuxin
Liu, Yang
Huang, Wenbing
author_facet Zhang, Yuelin
Cheng, Sijie
Li, Chen
Li, Zongzhao
Huang, Yuxin
Liu, Yang
Huang, Wenbing
contents Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning potential. Furthermore, processing long video trajectories with VLMs is computationally prohibitive for real-world deployment. To address these challenges, we propose the Recurrent Reasoning Vision-Language Model ($\text{R}^2$VLM). Our model features a recurrent reasoning framework that processes local video snippets iteratively, maintaining a global context through an evolving Chain of Thought (CoT). This CoT explicitly records task decomposition, key steps, and their completion status, enabling the model to reason about complex temporal dependencies. This design avoids the high cost of processing long videos while preserving essential reasoning capabilities. We train $\text{R}^2$VLM on large-scale, automatically generated datasets from ALFRED and Ego4D. Extensive experiments on progress estimation and downstream applications, including progress-enhanced policy learning, reward modeling for reinforcement learning, and proactive assistance, demonstrate that $\text{R}^2$VLM achieves strong performance and generalization, achieving a new state-of-the-art in long-horizon task progress estimation. The models and benchmarks are publicly available at \href{https://huggingface.co/collections/zhangyuelin/r2vlm}{huggingface}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
Zhang, Yuelin
Cheng, Sijie
Li, Chen
Li, Zongzhao
Huang, Yuxin
Liu, Yang
Huang, Wenbing
Computer Vision and Pattern Recognition
Artificial Intelligence
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning potential. Furthermore, processing long video trajectories with VLMs is computationally prohibitive for real-world deployment. To address these challenges, we propose the Recurrent Reasoning Vision-Language Model ($\text{R}^2$VLM). Our model features a recurrent reasoning framework that processes local video snippets iteratively, maintaining a global context through an evolving Chain of Thought (CoT). This CoT explicitly records task decomposition, key steps, and their completion status, enabling the model to reason about complex temporal dependencies. This design avoids the high cost of processing long videos while preserving essential reasoning capabilities. We train $\text{R}^2$VLM on large-scale, automatically generated datasets from ALFRED and Ego4D. Extensive experiments on progress estimation and downstream applications, including progress-enhanced policy learning, reward modeling for reinforcement learning, and proactive assistance, demonstrate that $\text{R}^2$VLM achieves strong performance and generalization, achieving a new state-of-the-art in long-horizon task progress estimation. The models and benchmarks are publicly available at \href{https://huggingface.co/collections/zhangyuelin/r2vlm}{huggingface}.
title Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.17312