E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Zhihao, Zhou, Jiaying, Zhang, Likui, Lv, Qinhan, Liu, Hao, Zhang, Jusheng, Li, Weizheng, Chen, Ziliang, Chen, Tianshui, Zhai, Ruifeng, Wang, Keze, Lin, Liang, Wang, Guangrun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910071683809280
author Zhan, Zhihao
Zhou, Jiaying
Zhang, Likui
Lv, Qinhan
Liu, Hao
Zhang, Jusheng
Li, Weizheng
Chen, Ziliang
Chen, Tianshui
Zhai, Ruifeng
Wang, Keze
Lin, Liang
Wang, Guangrun
author_facet Zhan, Zhihao
Zhou, Jiaying
Zhang, Likui
Lv, Qinhan
Liu, Hao
Zhang, Jusheng
Li, Weizheng
Chen, Ziliang
Chen, Tianshui
Zhai, Ruifeng
Wang, Keze
Lin, Liang
Wang, Guangrun
contents Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. We further introduce a spherical viewpoint perturbation augmentation to enhance robustness to camera shifts without additional data. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across 14 diverse environments, outperforming strong baselines by 10.7% on average.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
Zhan, Zhihao
Zhou, Jiaying
Zhang, Likui
Lv, Qinhan
Liu, Hao
Zhang, Jusheng
Li, Weizheng
Chen, Ziliang
Chen, Tianshui
Zhai, Ruifeng
Wang, Keze
Lin, Liang
Wang, Guangrun
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. We further introduce a spherical viewpoint perturbation augmentation to enhance robustness to camera shifts without additional data. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across 14 diverse environments, outperforming strong baselines by 10.7% on average.
title E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2511.21542