Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
Fuente:
arXiv
Saved in:
| Main Authors: | Niu, Jiahui, Gu, Kefan, Zhao, Yucheng, Liang, Shengwen, Wang, Tiancai, Hu, Xing, Wang, Ying, Li, Huawei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLaDA-VLA: Vision Language Diffusion Action Models
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
by: Chen, Yandu, et al.
Published: (2025)
by: Chen, Yandu, et al.
Published: (2025)
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
Running VLAs at Real-time Speed
by: Ma, Yunchao, et al.
Published: (2025)
by: Ma, Yunchao, et al.
Published: (2025)
ManiAgent: An Agentic Framework for General Robotic Manipulation
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
by: Wang, Siting, et al.
Published: (2026)
by: Wang, Siting, et al.
Published: (2026)
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
by: Song, Wenxuan, et al.
Published: (2026)
by: Song, Wenxuan, et al.
Published: (2026)
FASTER: Rethinking Real-Time Flow VLAs
by: Lu, Yuxiang, et al.
Published: (2026)
by: Lu, Yuxiang, et al.
Published: (2026)
cVLA: Towards Efficient Camera-Space VLAs
by: Argus, Max, et al.
Published: (2025)
by: Argus, Max, et al.
Published: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
by: Goyal, Ankit, et al.
Published: (2025)
by: Goyal, Ankit, et al.
Published: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
SITCOM: Scaling Inference-Time COMpute for VLAs
by: Saxena, Ayudh, et al.
Published: (2025)
by: Saxena, Ayudh, et al.
Published: (2025)
Towards foundational LiDAR world models with efficient latent flow matching
by: Liu, Tianran, et al.
Published: (2025)
by: Liu, Tianran, et al.
Published: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
by: Fang, Yu, et al.
Published: (2026)
by: Fang, Yu, et al.
Published: (2026)
PoseINN: Realtime Visual-based Pose Regression and Localization with Invertible Neural Networks
by: Zang, Zirui, et al.
Published: (2024)
by: Zang, Zirui, et al.
Published: (2024)
SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
by: Ni, Chaojun, et al.
Published: (2025)
by: Ni, Chaojun, et al.
Published: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
Realtime Robust Shape Estimation of Deformable Linear Object
by: Zhang, Jiaming, et al.
Published: (2024)
by: Zhang, Jiaming, et al.
Published: (2024)
A Pragmatic VLA Foundation Model
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
FLASH: Efficient Visuomotor Policy via Sparse Sampling
by: Bai, Jiaqi, et al.
Published: (2026)
by: Bai, Jiaqi, et al.
Published: (2026)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
by: Huang, Binyuan, et al.
Published: (2024)
by: Huang, Binyuan, et al.
Published: (2024)
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
by: Guo, Xinyu, et al.
Published: (2026)
by: Guo, Xinyu, et al.
Published: (2026)
GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
by: Neau, Maëlic, et al.
Published: (2025)
by: Neau, Maëlic, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
by: Guo, Heyu, et al.
Published: (2025)
by: Guo, Heyu, et al.
Published: (2025)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
Shallow-π: Knowledge Distillation for Flow-based VLAs
by: Jeon, Boseong, et al.
Published: (2026)
by: Jeon, Boseong, et al.
Published: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
by: Liang, Zhixuan, et al.
Published: (2025)
by: Liang, Zhixuan, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
by: Ma, Chuanhao, et al.
Published: (2026)
by: Ma, Chuanhao, et al.
Published: (2026)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
by: Liu, Chuhang, et al.
Published: (2026)
by: Liu, Chuhang, et al.
Published: (2026)
Similar Items
-
LLaDA-VLA: Vision Language Diffusion Action Models
by: Wen, Yuqing, et al.
Published: (2025) -
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
by: Chen, Yandu, et al.
Published: (2025) -
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
by: Yang, Chen, et al.
Published: (2026) -
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025) -
Running VLAs at Real-time Speed
by: Ma, Yunchao, et al.
Published: (2025)