StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Yiran, Guo, Dongqi, Zhao, Tianchen, Gao, Feng, Shi, Liangzhi, Yu, Chao, Mo, ZhiJian, Xiao, Qihua, Peng, XiaoShuai, Liao, Qingmin, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
von: Fang, Zhou, et al.
Veröffentlicht: (2026)
von: Fang, Zhou, et al.
Veröffentlicht: (2026)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
von: Pan, Xu, et al.
Veröffentlicht: (2026)
von: Pan, Xu, et al.
Veröffentlicht: (2026)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
von: Zang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zang, Hongzhi, et al.
Veröffentlicht: (2025)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
von: Kang, Hyolim, et al.
Veröffentlicht: (2024)
von: Kang, Hyolim, et al.
Veröffentlicht: (2024)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
von: Yu, Wenda, et al.
Veröffentlicht: (2026)
von: Yu, Wenda, et al.
Veröffentlicht: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
EdgeVLA: Efficient Vision-Language-Action Models
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
Action-to-Action Flow Matching
von: Jia, Jindou, et al.
Veröffentlicht: (2026)
von: Jia, Jindou, et al.
Veröffentlicht: (2026)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
von: Yan, Yibin, et al.
Veröffentlicht: (2026)
von: Yan, Yibin, et al.
Veröffentlicht: (2026)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
VITA: Vision-to-Action Flow Matching Policy
von: Gao, Dechen, et al.
Veröffentlicht: (2025)
von: Gao, Dechen, et al.
Veröffentlicht: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
von: Wang, Si-Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Si-Cheng, et al.
Veröffentlicht: (2025)
Action Emergence from Streaming Intent
von: Jing, Pengfei, et al.
Veröffentlicht: (2026)
von: Jing, Pengfei, et al.
Veröffentlicht: (2026)
Dual-Stream Alignment for Action Segmentation
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
von: Liu, Zhi
Veröffentlicht: (2026)
von: Liu, Zhi
Veröffentlicht: (2026)
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio
von: Jiang, Yujiao, et al.
Veröffentlicht: (2026)
von: Jiang, Yujiao, et al.
Veröffentlicht: (2026)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
OpenVLA: An Open-Source Vision-Language-Action Model
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
von: Arai, Hidehisa, et al.
Veröffentlicht: (2024)
von: Arai, Hidehisa, et al.
Veröffentlicht: (2024)
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
von: Hu, Yutong, et al.
Veröffentlicht: (2026)
von: Hu, Yutong, et al.
Veröffentlicht: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action
von: Shi, Xiang, et al.
Veröffentlicht: (2026)
von: Shi, Xiang, et al.
Veröffentlicht: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025)
von: Won, John, et al.
Veröffentlicht: (2025)
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
von: Lee, Jimin, et al.
Veröffentlicht: (2026)
von: Lee, Jimin, et al.
Veröffentlicht: (2026)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
von: Fang, Zhou, et al.
Veröffentlicht: (2026) -
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
von: Pan, Xu, et al.
Veröffentlicht: (2026) -
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025) -
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
von: Zang, Hongzhi, et al.
Veröffentlicht: (2025) -
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)