ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chi-Pin, Wu, Yueh-Hua, Chen, Min-Hung, Wang, Yu-Chiang Frank, Yang, Fu-En
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916955980562432
author Huang, Chi-Pin
Wu, Yueh-Hua
Chen, Min-Hung
Wang, Yu-Chiang Frank
Yang, Fu-En
author_facet Huang, Chi-Pin
Wu, Yueh-Hua
Chen, Min-Hung
Wang, Yu-Chiang Frank
Yang, Fu-En
contents Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Huang, Chi-Pin
Wu, Yueh-Hua
Chen, Min-Hung
Wang, Yu-Chiang Frank
Yang, Fu-En
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks.
title ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2507.16815