From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yajie, Zhang, Bozhou, Gu, Chun, Ma, Zipei, Zhang, Jiahui, Deng, Jiankang, Zhu, Xiatian, Zhang, Li
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916005320589312
author Li, Yajie
Zhang, Bozhou
Gu, Chun
Ma, Zipei
Zhang, Jiahui
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
author_facet Li, Yajie
Zhang, Bozhou
Gu, Chun
Ma, Zipei
Zhang, Jiahui
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
contents Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or directly decode generated videos into actions, both suffering from a mismatch between visual realism and control relevance. As a result, predicted observations emphasize perceptual fidelity rather than action-centric causes of state transitions, leading to indirect and unstable control. To address this gap, we propose MoLA (Mixture of Latent Actions), a control-oriented interface that transforms imagined future videos into executable representations. Instead of passing predicted frames directly to the policy, MoLA leverages a mixture of pretrained inverse dynamics models to infer a mixture of latent actions implied by generated visual transitions. These modality-aware inverse dynamics models capture complementary semantic, depth, and flow cues, providing a structured and physically grounded action representation that bridges video imagination and policy execution. We evaluate our approach on simulated benchmarks (LIBERO, CALVIN, and LIBERO-Plus) and real-world robot manipulation tasks, achieving consistent gains in task success, temporal consistency, and generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12167
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
Li, Yajie
Zhang, Bozhou
Gu, Chun
Ma, Zipei
Zhang, Jiahui
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
Robotics
Computer Vision and Pattern Recognition
Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or directly decode generated videos into actions, both suffering from a mismatch between visual realism and control relevance. As a result, predicted observations emphasize perceptual fidelity rather than action-centric causes of state transitions, leading to indirect and unstable control. To address this gap, we propose MoLA (Mixture of Latent Actions), a control-oriented interface that transforms imagined future videos into executable representations. Instead of passing predicted frames directly to the policy, MoLA leverages a mixture of pretrained inverse dynamics models to infer a mixture of latent actions implied by generated visual transitions. These modality-aware inverse dynamics models capture complementary semantic, depth, and flow cues, providing a structured and physically grounded action representation that bridges video imagination and policy execution. We evaluate our approach on simulated benchmarks (LIBERO, CALVIN, and LIBERO-Plus) and real-world robot manipulation tasks, achieving consistent gains in task success, temporal consistency, and generalization.
title From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12167