NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Zheng, Liu, Mingyu, Lin, Xiaoyi, Zhu, Muzhi, Zhao, Canyu, Du, Zongze, Lin, Ye, Li, Xiaoman, Jia, Yiduo, Zhong, Hao, Chen, Hao, Shen, Chunhua
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909032230420480
author Huang, Zheng
Liu, Mingyu
Lin, Xiaoyi
Zhu, Muzhi
Zhao, Canyu
Du, Zongze
Lin, Ye
Li, Xiaoman
Jia, Yiduo
Zhong, Hao
Chen, Hao
Shen, Chunhua
author_facet Huang, Zheng
Liu, Mingyu
Lin, Xiaoyi
Zhu, Muzhi
Zhao, Canyu
Du, Zongze
Lin, Ye
Li, Xiaoman
Jia, Yiduo
Zhong, Hao
Chen, Hao
Shen, Chunhua
contents Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic forgetting. This issue stems from their overreliance on continuous action sequences or action chunks, which inadvertently create isolated data silos that disrupt knowledge retention across tasks. To tackle these challenges, we propose the Narrowing of Trajectory VLA (NoTVLA) framework: a novel approach that narrows its focus to sparse trajectories, thereby avoiding the catastrophic forgetting associated with dense trajectory fine-tuning. A key innovation of NoTVLA lies in its trajectory planning strategy: instead of centering on the target object's trajectory, it leverages temporal compression and spatial reasoning pruning specifically for the robot end effector's trajectory. Furthermore, training is conducted using these sparse trajectories rather than dense action trajectories, an optimization that delivers remarkable practical advantages with better performance in zero-shot. In multi-task evaluation scenarios, NoTVLA achieves superior performance and generalization compared to pi0 while operating under two critical constraints: it uses over an order of magnitude less computing power than pi0 and requires no wrist-mounted camera. This design ensures that NoTVLA's operational accuracy closely approximates that of single-task expert models. Crucially, it also preserves the model's inherent language capabilities, enabling zero-shot generalization in specific scenarios, supporting unified model deployment across multiple robot platforms, and fostering a degree of generalization even when perceiving tasks from novel perspectives.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03895
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Huang, Zheng
Liu, Mingyu
Lin, Xiaoyi
Zhu, Muzhi
Zhao, Canyu
Du, Zongze
Lin, Ye
Li, Xiaoman
Jia, Yiduo
Zhong, Hao
Chen, Hao
Shen, Chunhua
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic forgetting. This issue stems from their overreliance on continuous action sequences or action chunks, which inadvertently create isolated data silos that disrupt knowledge retention across tasks. To tackle these challenges, we propose the Narrowing of Trajectory VLA (NoTVLA) framework: a novel approach that narrows its focus to sparse trajectories, thereby avoiding the catastrophic forgetting associated with dense trajectory fine-tuning. A key innovation of NoTVLA lies in its trajectory planning strategy: instead of centering on the target object's trajectory, it leverages temporal compression and spatial reasoning pruning specifically for the robot end effector's trajectory. Furthermore, training is conducted using these sparse trajectories rather than dense action trajectories, an optimization that delivers remarkable practical advantages with better performance in zero-shot. In multi-task evaluation scenarios, NoTVLA achieves superior performance and generalization compared to pi0 while operating under two critical constraints: it uses over an order of magnitude less computing power than pi0 and requires no wrist-mounted camera. This design ensures that NoTVLA's operational accuracy closely approximates that of single-task expert models. Crucially, it also preserves the model's inherent language capabilities, enabling zero-shot generalization in specific scenarios, supporting unified model deployment across multiple robot platforms, and fostering a degree of generalization even when perceiving tasks from novel perspectives.
title NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03895