Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917352329707520 |
|---|---|
| author | Zhao, Chen Wang, Zhuoran Li, Haoyang Bao, Shifeng Li, Guanlin Feng, Youhe Li, Yang Tang, Jie Zhang, Jing |
| author_facet | Zhao, Chen Wang, Zhuoran Li, Haoyang Bao, Shifeng Li, Guanlin Feng, Youhe Li, Yang Tang, Jie Zhang, Jing |
| contents | Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generate high-precision continuous action chunks, while auto-regressive generation can be slower and less accurate at low-level control. Yet auto-regressive paradigms still provide complementary priors that can improve robustness and generalization in out-of-distribution environments. To leverage both paradigms, we propose Action-Draft-and-Verify (ADV): diffusion action expert drafts multiple candidate action chunks, and the VLM selects one by scoring all candidates in a single forward pass with a perplexity-style metric. Under matched backbones, training data, and action-chunk length, ADV improves success rate by +4.3 points in simulation and +19.7 points in real-world over diffusion-based baseline, with a single-pass VLM reranking overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_18091 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model Zhao, Chen Wang, Zhuoran Li, Haoyang Bao, Shifeng Li, Guanlin Feng, Youhe Li, Yang Tang, Jie Zhang, Jing Computer Vision and Pattern Recognition Robotics Vision-Language-Action (VLA) models have recently demonstrated strong performance across embodied tasks. Modern VLAs commonly employ diffusion action experts to efficiently generate high-precision continuous action chunks, while auto-regressive generation can be slower and less accurate at low-level control. Yet auto-regressive paradigms still provide complementary priors that can improve robustness and generalization in out-of-distribution environments. To leverage both paradigms, we propose Action-Draft-and-Verify (ADV): diffusion action expert drafts multiple candidate action chunks, and the VLM selects one by scoring all candidates in a single forward pass with a perplexity-style metric. Under matched backbones, training data, and action-chunk length, ADV improves success rate by +4.3 points in simulation and +19.7 points in real-world over diffusion-based baseline, with a single-pass VLM reranking overhead. |
| title | Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2603.18091 |