KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Gaoge, Gao, Zhengqing, Li, Ziwen, Huang, Jiaxin, Huang, Shaoli, Karray, Fakhri, Gong, Mingming, Liu, Tongliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917351465680896
author Han, Gaoge
Gao, Zhengqing
Li, Ziwen
Huang, Jiaxin
Huang, Shaoli
Karray, Fakhri
Gong, Mingming
Liu, Tongliang
author_facet Han, Gaoge
Gao, Zhengqing
Li, Ziwen
Huang, Jiaxin
Huang, Shaoli
Karray, Fakhri
Gong, Mingming
Liu, Tongliang
contents In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, trajectory, orientation, and relative displacement) from initiation through completion, at key moments, unlike existing action instructions that capture kinematics only coarsely or partially, thereby supporting fine-grained and personalized manipulation. In this setting, where task goals remain invariant while execution trajectories must adapt to instruction-level kinematic specifications. To address this challenge, we propose KineVLA, a vision-language-action framework that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve as explicit, supervised intermediate variables that align language and action. To support this task, we construct the kinematics-aware VLA datasets spanning both simulation and real-world robotic platforms, featuring instruction-level kinematic variations and bi-level annotations. Extensive experiments on LIBERO and a Realman-75 robot demonstrate that KineVLA consistently outperforms strong VLA baselines on kinematics-sensitive benchmarks, achieving more precise, controllable, and generalizable manipulation behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17524
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
Han, Gaoge
Gao, Zhengqing
Li, Ziwen
Huang, Jiaxin
Huang, Shaoli
Karray, Fakhri
Gong, Mingming
Liu, Tongliang
Robotics
Artificial Intelligence
In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, trajectory, orientation, and relative displacement) from initiation through completion, at key moments, unlike existing action instructions that capture kinematics only coarsely or partially, thereby supporting fine-grained and personalized manipulation. In this setting, where task goals remain invariant while execution trajectories must adapt to instruction-level kinematic specifications. To address this challenge, we propose KineVLA, a vision-language-action framework that explicitly decouples goal-level invariance from kinematics-level variability through a bi-level action representation and bi-level reasoning tokens to serve as explicit, supervised intermediate variables that align language and action. To support this task, we construct the kinematics-aware VLA datasets spanning both simulation and real-world robotic platforms, featuring instruction-level kinematic variations and bi-level annotations. Extensive experiments on LIBERO and a Realman-75 robot demonstrate that KineVLA consistently outperforms strong VLA baselines on kinematics-sensitive benchmarks, achieving more precise, controllable, and generalizable manipulation behaviors.
title KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2603.17524