Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Chen, Wang, Zhuoran, Li, Haoyang, Bao, Shifeng, Li, Guanlin, Feng, Youhe, Li, Yang, Tang, Jie, Zhang, Jing
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917352329707520
author Zhao, Chen
Wang, Zhuoran
Li, Haoyang
Bao, Shifeng
Li, Guanlin
Feng, Youhe
Li, Yang
Tang, Jie
Zhang, Jing
author_facet Zhao, Chen
Wang, Zhuoran
Li, Haoyang
Bao, Shifeng
Li, Guanlin
Feng, Youhe
Li, Yang
Tang, Jie
Zhang, Jing
contents Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generate high-precision continuous action chunks, while auto-regressive generation can be slower and less accurate at low-level control. Yet auto-regressive paradigms still provide complementary priors that can improve robustness and generalization in out-of-distribution environments. To leverage both paradigms, we propose Action-Draft-and-Verify (ADV): diffusion action expert drafts multiple candidate action chunks, and the VLM selects one by scoring all candidates in a single forward pass with a perplexity-style metric. Under matched backbones, training data, and action-chunk length, ADV improves success rate by +4.3 points in simulation and +19.7 points in real-world over diffusion-based baseline, with a single-pass VLM reranking overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18091
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Zhao, Chen
Wang, Zhuoran
Li, Haoyang
Bao, Shifeng
Li, Guanlin
Feng, Youhe
Li, Yang
Tang, Jie
Zhang, Jing
Computer Vision and Pattern Recognition
Robotics
Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generate high-precision continuous action chunks, while auto-regressive generation can be slower and less accurate at low-level control. Yet auto-regressive paradigms still provide complementary priors that can improve robustness and generalization in out-of-distribution environments. To leverage both paradigms, we propose Action-Draft-and-Verify (ADV): diffusion action expert drafts multiple candidate action chunks, and the VLM selects one by scoring all candidates in a single forward pass with a perplexity-style metric. Under matched backbones, training data, and action-chunk length, ADV improves success rate by +4.3 points in simulation and +19.7 points in real-world over diffusion-based baseline, with a single-pass VLM reranking overhead.
title Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2603.18091