DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Jiayi, Song, Wenxuan, Chen, Shuai, Wang, Jingbo, Li, Zhijun, Li, Haoang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
von: Wang, Sen, et al.
Veröffentlicht: (2025)
von: Wang, Sen, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
KAN We Flow? Advancing Robotic Manipulation with 3D Flow Matching via KAN & RWKV
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
von: Jiang, Yuming, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
von: Chopra, Samarth, et al.
Veröffentlicht: (2025)
von: Chopra, Samarth, et al.
Veröffentlicht: (2025)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
von: Ma, Chuanhao, et al.
Veröffentlicht: (2026)
von: Ma, Chuanhao, et al.
Veröffentlicht: (2026)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
von: Cui, Can, et al.
Veröffentlicht: (2025)
von: Cui, Can, et al.
Veröffentlicht: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
Demystifying Action Space Design for Robotic Manipulation Policies
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
von: Li, Ziwen, et al.
Veröffentlicht: (2025)
von: Li, Ziwen, et al.
Veröffentlicht: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
von: Li, Yajie, et al.
Veröffentlicht: (2026)
von: Li, Yajie, et al.
Veröffentlicht: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
von: Zhao, Guoyang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
von: Song, Wenxuan, et al.
Veröffentlicht: (2026) -
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
von: Chen, Jiayi, et al.
Veröffentlicht: (2025) -
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025) -
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025) -
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)