HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiaming, Chen, Hao, An, Pengju, Liu, Zhuoyang, Zhang, Renrui, Gu, Chenyang, Li, Xiaoqi, Guo, Ziyu, Chen, Sixiang, Liu, Mengzhen, Hou, Chengkai, Zhao, Mengdi, Zhou, KC alex, Heng, Pheng-Ann, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
by: Liu, Zhuoyang, et al.
Published: (2026)
by: Liu, Zhuoyang, et al.
Published: (2026)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
by: Li, Chenxuan, et al.
Published: (2024)
by: Li, Chenxuan, et al.
Published: (2024)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
by: Fang, Zhen, et al.
Published: (2025)
by: Fang, Zhen, et al.
Published: (2025)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
by: Luo, Yulin, et al.
Published: (2026)
by: Luo, Yulin, et al.
Published: (2026)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
by: Jia, Yueru, et al.
Published: (2024)
by: Jia, Yueru, et al.
Published: (2024)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
by: Li, Guangrun, et al.
Published: (2025)
by: Li, Guangrun, et al.
Published: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
by: Wen, Junjie, et al.
Published: (2025)
by: Wen, Junjie, et al.
Published: (2025)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
by: Pan, Mingjie, et al.
Published: (2023)
by: Pan, Mingjie, et al.
Published: (2023)
Swiss Air Cancellation Policy
by: alex
Published: (2025)
by: alex
Published: (2025)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
by: Guo, Ziyu, et al.
Published: (2024)
by: Guo, Ziyu, et al.
Published: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
by: Liu, Jiaming, et al.
Published: (2023)
by: Liu, Jiaming, et al.
Published: (2023)
Self-Organized Criticality in Meta-Cognitive Control Loops of Large Language Models
by: smith, alex
Published: (2025)
by: smith, alex
Published: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
by: Tong, Chengzhuo, et al.
Published: (2025)
by: Tong, Chengzhuo, et al.
Published: (2025)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
by: Wang, Jinghao, et al.
Published: (2026)
by: Wang, Jinghao, et al.
Published: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
by: Cao, Jiajun, et al.
Published: (2026)
by: Cao, Jiajun, et al.
Published: (2026)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
by: Liu, Jiaming, et al.
Published: (2023)
by: Liu, Jiaming, et al.
Published: (2023)
WorldVLA: Towards Autoregressive Action World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection
by: Liu, Jiaming, et al.
Published: (2022)
by: Liu, Jiaming, et al.
Published: (2022)
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
by: Chen, Xinyan, et al.
Published: (2023)
by: Chen, Xinyan, et al.
Published: (2023)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning
by: Wang, Ruiheng, et al.
Published: (2026)
by: Wang, Ruiheng, et al.
Published: (2026)
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation
by: Ni, Jiayi, et al.
Published: (2023)
by: Ni, Jiayi, et al.
Published: (2023)
Similar Items
-
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025) -
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
by: Liu, Zhuoyang, et al.
Published: (2026) -
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
by: Liu, Jiaming, et al.
Published: (2024) -
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
by: Chen, Hao, et al.
Published: (2025) -
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
by: Chen, Sixiang, et al.
Published: (2025)