ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Minho, Kim, Kinam, Hyung, Junha, Jang, Hyojin, Jin, Hoiyeong, Yun, Jooyeol, Lee, Hojoon, Choo, Jaegul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914420788035584
author Park, Minho
Kim, Kinam
Hyung, Junha
Jang, Hyojin
Jin, Hoiyeong
Yun, Jooyeol
Lee, Hojoon
Choo, Jaegul
author_facet Park, Minho
Kim, Kinam
Hyung, Junha
Jang, Hyojin
Jin, Hoiyeong
Yun, Jooyeol
Lee, Hojoon
Choo, Jaegul
contents Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstrations: jerks, pauses, and jitter which reduce action coherence. Reduced action coherence causes instability and trajectory drift during deployment, failures that are catastrophic in fine-grained manipulation where precision is crucial. In this paper, we present Action Coherence Guidance (ACG) for VLA models, a training-free test-time guidance algorithm that improves action coherence and thereby yields performance gains. Evaluated on RoboCasa, DexMimicGen, and real-world SO-101 tasks, ACG consistently improves action coherence and boosts success rates across diverse manipulation tasks. Code and project page are available at https://github.com/DAVIAN-Robotics/ACG and https://DAVIAN-Robotics.github.io/ACG , respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22201
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Park, Minho
Kim, Kinam
Hyung, Junha
Jang, Hyojin
Jin, Hoiyeong
Yun, Jooyeol
Lee, Hojoon
Choo, Jaegul
Robotics
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstrations: jerks, pauses, and jitter which reduce action coherence. Reduced action coherence causes instability and trajectory drift during deployment, failures that are catastrophic in fine-grained manipulation where precision is crucial. In this paper, we present Action Coherence Guidance (ACG) for VLA models, a training-free test-time guidance algorithm that improves action coherence and thereby yields performance gains. Evaluated on RoboCasa, DexMimicGen, and real-world SO-101 tasks, ACG consistently improves action coherence and boosts success rates across diverse manipulation tasks. Code and project page are available at https://github.com/DAVIAN-Robotics/ACG and https://DAVIAN-Robotics.github.io/ACG , respectively.
title ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
topic Robotics
url https://arxiv.org/abs/2510.22201