ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Minho, Kim, Kinam, Hyung, Junha, Jang, Hyojin, Jin, Hoiyeong, Yun, Jooyeol, Lee, Hojoon, Choo, Jaegul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025)
by: Kim, Kinam, et al.
Published: (2025)
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
by: Hwang, Sungwon, et al.
Published: (2025)
by: Hwang, Sungwon, et al.
Published: (2025)
Regularized Training with Generated Datasets for Name-Only Transfer of Vision-Language Models
by: Park, Minho, et al.
Published: (2024)
by: Park, Minho, et al.
Published: (2024)
PHUMA: Physically-Grounded Humanoid Locomotion Dataset
by: Lee, Kyungmin, et al.
Published: (2025)
by: Lee, Kyungmin, et al.
Published: (2025)
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
by: Kim, Donghu, et al.
Published: (2026)
by: Kim, Donghu, et al.
Published: (2026)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
by: Yun, Jooyeol, et al.
Published: (2025)
by: Yun, Jooyeol, et al.
Published: (2025)
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
by: Jo, Kyungmin, et al.
Published: (2025)
by: Jo, Kyungmin, et al.
Published: (2025)
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Enabling Region-Specific Control via Lassos in Point-Based Colorization
by: Lee, Sanghyeon, et al.
Published: (2024)
by: Lee, Sanghyeon, et al.
Published: (2024)
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
by: Lee, Jimin, et al.
Published: (2026)
by: Lee, Jimin, et al.
Published: (2026)
MagiCapture: High-Resolution Multi-Concept Portrait Customization
by: Hyung, Junha, et al.
Published: (2023)
by: Hyung, Junha, et al.
Published: (2023)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
Learning to Act Robustly with View-Invariant Latent Actions
by: Jeong, Youngjoon, et al.
Published: (2026)
by: Jeong, Youngjoon, et al.
Published: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
by: Park, Jongwoo, et al.
Published: (2026)
by: Park, Jongwoo, et al.
Published: (2026)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
by: Park, Jeongeun, et al.
Published: (2025)
by: Park, Jeongeun, et al.
Published: (2025)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
by: Won, John, et al.
Published: (2025)
by: Won, John, et al.
Published: (2025)
Verifier-free Test-Time Sampling for Vision Language Action Models
by: Jang, Suhyeok, et al.
Published: (2025)
by: Jang, Suhyeok, et al.
Published: (2025)
Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
by: Gwak, Daehoon, et al.
Published: (2025)
by: Gwak, Daehoon, et al.
Published: (2025)
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
by: Gwak, Daehoon, et al.
Published: (2026)
by: Gwak, Daehoon, et al.
Published: (2026)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
by: Dai, Yuntao, et al.
Published: (2025)
by: Dai, Yuntao, et al.
Published: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
by: Liang, Yuanchang, et al.
Published: (2026)
by: Liang, Yuanchang, et al.
Published: (2026)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
by: Hwang, Dongyoon, et al.
Published: (2025)
by: Hwang, Dongyoon, et al.
Published: (2025)
Adaptive Capacity Allocation for Vision Language Action Fine-tuning
by: Kim, Donghoon, et al.
Published: (2026)
by: Kim, Donghoon, et al.
Published: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation
by: Jung, Kyudan, et al.
Published: (2025)
by: Jung, Kyudan, et al.
Published: (2025)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
by: Fang, Zhou, et al.
Published: (2026)
by: Fang, Zhou, et al.
Published: (2026)
Mean-Flow based One-Step Vision-Language-Action
by: Chen, Yang, et al.
Published: (2026)
by: Chen, Yang, et al.
Published: (2026)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
by: Heo, Hyeongjun, et al.
Published: (2026)
by: Heo, Hyeongjun, et al.
Published: (2026)
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
by: Kim, Sungha, et al.
Published: (2026)
by: Kim, Sungha, et al.
Published: (2026)
Similar Items
-
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025) -
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025) -
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
by: Hwang, Sungwon, et al.
Published: (2025) -
Regularized Training with Generated Datasets for Name-Only Transfer of Vision-Language Models
by: Park, Minho, et al.
Published: (2024) -
PHUMA: Physically-Grounded Humanoid Locomotion Dataset
by: Lee, Kyungmin, et al.
Published: (2025)