ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Yuntao, Gu, Hang, Wang, Teng, Cheng, Qianyu, Zheng, Yifei, Qiu, Zhiyong, Gong, Lei, Lou, Wenqi, Zhou, Xuehai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActionFlow: Equivariant, Accurate, and Efficient Policies with Spatially Symmetric Flow Matching
by: Funk, Niklas, et al.
Published: (2024)
by: Funk, Niklas, et al.
Published: (2024)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
EdgeVLA: Efficient Vision-Language-Action Models
by: Budzianowski, Paweł, et al.
Published: (2025)
by: Budzianowski, Paweł, et al.
Published: (2025)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
by: Liu, Yicheng, et al.
Published: (2025)
by: Liu, Yicheng, et al.
Published: (2025)
Toward Embodiment Equivariant Vision-Language-Action Policy
by: Chen, Anzhe, et al.
Published: (2025)
by: Chen, Anzhe, et al.
Published: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2025)
by: Zhong, Zhide, et al.
Published: (2025)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
by: Zhou, Kaijun, et al.
Published: (2026)
by: Zhou, Kaijun, et al.
Published: (2026)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
by: Shen, Weijie, et al.
Published: (2025)
by: Shen, Weijie, et al.
Published: (2025)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
by: Fang, Zhou, et al.
Published: (2026)
by: Fang, Zhou, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
Mean-Flow based One-Step Vision-Language-Action
by: Chen, Yang, et al.
Published: (2026)
by: Chen, Yang, et al.
Published: (2026)
Dexbotic: Open-Source Vision-Language-Action Toolbox
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Action Hallucination in Generative Vision-Language-Action Models
by: Soh, Harold, et al.
Published: (2026)
by: Soh, Harold, et al.
Published: (2026)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
by: Liang, Yuanchang, et al.
Published: (2026)
by: Liang, Yuanchang, et al.
Published: (2026)
Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
by: Niu, Haochen, et al.
Published: (2026)
by: Niu, Haochen, et al.
Published: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026)
by: Dong, Zibin, et al.
Published: (2026)
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
by: Qiu, Weikang, et al.
Published: (2026)
by: Qiu, Weikang, et al.
Published: (2026)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
by: Wulff, Theodor, et al.
Published: (2026)
by: Wulff, Theodor, et al.
Published: (2026)
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
by: Liu, Haoyun, et al.
Published: (2026)
by: Liu, Haoyun, et al.
Published: (2026)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
by: Deng, Shengliang, et al.
Published: (2025)
by: Deng, Shengliang, et al.
Published: (2025)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
by: Pan, Xu, et al.
Published: (2026)
by: Pan, Xu, et al.
Published: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
by: Li, Zuolei, et al.
Published: (2025)
by: Li, Zuolei, et al.
Published: (2025)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model
by: Li, Shunlei, et al.
Published: (2024)
by: Li, Shunlei, et al.
Published: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
by: Gu, Hang, et al.
Published: (2026)
by: Gu, Hang, et al.
Published: (2026)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Similar Items
-
ActionFlow: Equivariant, Accurate, and Efficient Policies with Spatially Symmetric Flow Matching
by: Funk, Niklas, et al.
Published: (2024) -
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026) -
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
by: Park, Minho, et al.
Published: (2025) -
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026) -
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)