Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yudong, Li, Yuan, Tang, Zijia, Zheng, Yuxi, Lin, Yueqian, Wang, Qinsi, Li, Yi, Liu, Shuangjun, Zhang, Shuai, Jing, Taotao, Gao, Dashan, Bi, Ning, Sun, Jingwei, Chen, Yiran, Li, Hai
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918481526521856
author Liu, Yudong
Li, Yuan
Tang, Zijia
Zheng, Yuxi
Lin, Yueqian
Wang, Qinsi
Li, Yi
Liu, Shuangjun
Zhang, Shuai
Jing, Taotao
Gao, Dashan
Bi, Ning
Sun, Jingwei
Chen, Yiran
Li, Hai
author_facet Liu, Yudong
Li, Yuan
Tang, Zijia
Zheng, Yuxi
Lin, Yueqian
Wang, Qinsi
Li, Yi
Liu, Shuangjun
Zhang, Shuai
Jing, Taotao
Gao, Dashan
Bi, Ning
Sun, Jingwei
Chen, Yiran
Li, Hai
contents Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step while producing temporally redundant features. We propose Latent Bridge, a lightweight model that predicts VLM output deltas between timesteps, enabling the action head to operate on predicted outputs while the expensive VLM backbone is called only periodically. We instantiate Latent Bridge on two architecturally distinct VLAs: GR00T-N1.6 (feature-space bridge) and π0.5 (KV-cache bridge), demonstrating that the approach generalizes across VLA designs. Our task-agnostic DAgger training pipeline transfers across benchmarks without modification. Across four LIBERO suites, 24 RoboCasa kitchen tasks, and the ALOHA sim transfer-cube task, Latent Bridge achieves 95-100% performance retention while reducing VLM calls by 50-75%, yielding 1.65-1.73x net per-episode speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
Liu, Yudong
Li, Yuan
Tang, Zijia
Zheng, Yuxi
Lin, Yueqian
Wang, Qinsi
Li, Yi
Liu, Shuangjun
Zhang, Shuai
Jing, Taotao
Gao, Dashan
Bi, Ning
Sun, Jingwei
Chen, Yiran
Li, Hai
Robotics
Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step while producing temporally redundant features. We propose Latent Bridge, a lightweight model that predicts VLM output deltas between timesteps, enabling the action head to operate on predicted outputs while the expensive VLM backbone is called only periodically. We instantiate Latent Bridge on two architecturally distinct VLAs: GR00T-N1.6 (feature-space bridge) and π0.5 (KV-cache bridge), demonstrating that the approach generalizes across VLA designs. Our task-agnostic DAgger training pipeline transfers across benchmarks without modification. Across four LIBERO suites, 24 RoboCasa kitchen tasks, and the ALOHA sim transfer-cube task, Latent Bridge achieves 95-100% performance retention while reducing VLM calls by 50-75%, yielding 1.65-1.73x net per-episode speedup.
title Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
topic Robotics
url https://arxiv.org/abs/2605.02739