IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Shijie, Yu, Bin, Lin, Xiaopeng, Shen, Zhaolong, Yang, Laurence Tianruo, Jin, Yurun, Liu, Haishan, Wu, Changti, Yuan, Hang, Huang, Cong, Chen, Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913128964423680
author Lian, Shijie
Yu, Bin
Lin, Xiaopeng
Shen, Zhaolong
Yang, Laurence Tianruo
Jin, Yurun
Liu, Haishan
Wu, Changti
Yuan, Hang
Huang, Cong
Chen, Kai
author_facet Lian, Shijie
Yu, Bin
Lin, Xiaopeng
Shen, Zhaolong
Yang, Laurence Tianruo
Jin, Yurun
Liu, Haishan
Wu, Changti
Yuan, Hang
Huang, Cong
Chen, Kai
contents Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines
format Preprint
id arxiv_https___arxiv_org_abs_2605_14712
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
Lian, Shijie
Yu, Bin
Lin, Xiaopeng
Shen, Zhaolong
Yang, Laurence Tianruo
Jin, Yurun
Liu, Haishan
Wu, Changti
Yuan, Hang
Huang, Cong
Chen, Kai
Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines
title IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
topic Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.14712