ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Chenyang, Liu, Jiaming, Chen, Hao, Huang, Runzhong, Wuwu, Qingpo, Liu, Zhuoyang, Li, Xiaoqi, Li, Ying, Zhang, Renrui, Jia, Peng, Heng, Pheng-Ann, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2026)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Altered Thoughts, Altered Actions: Probing Chain-of-Thought Vulnerabilities in VLA Robotic Manipulation
von: Trinh, Tuan Duong, et al.
Veröffentlicht: (2026)
von: Trinh, Tuan Duong, et al.
Veröffentlicht: (2026)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
von: Zhou, Hui, et al.
Veröffentlicht: (2025)
von: Zhou, Hui, et al.
Veröffentlicht: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
RoboArmGS: High-Quality Robotic Arm Splatting via Bézier Curve Refinement
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
von: Jiang, Haoran, et al.
Veröffentlicht: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
von: Li, Yutai, et al.
Veröffentlicht: (2026)
von: Li, Yutai, et al.
Veröffentlicht: (2026)
ST-$π$: Structured SpatioTemporal VLA for Robotic Manipulation
von: Ma, Chuanhao, et al.
Veröffentlicht: (2026)
von: Ma, Chuanhao, et al.
Veröffentlicht: (2026)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
von: Lin, Min, et al.
Veröffentlicht: (2025)
von: Lin, Min, et al.
Veröffentlicht: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
von: Xu, Qinwen, et al.
Veröffentlicht: (2026)
von: Xu, Qinwen, et al.
Veröffentlicht: (2026)
DroneVLA: VLA based Aerial Manipulation
von: Mehboob, Fawad, et al.
Veröffentlicht: (2026)
von: Mehboob, Fawad, et al.
Veröffentlicht: (2026)
GazeVLA: Learning Human Intention for Robotic Manipulation
von: Li, Chengyang, et al.
Veröffentlicht: (2026)
von: Li, Chengyang, et al.
Veröffentlicht: (2026)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
von: Cao, Jiajun, et al.
Veröffentlicht: (2026)
von: Cao, Jiajun, et al.
Veröffentlicht: (2026)
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
von: Heng, Liang, et al.
Veröffentlicht: (2025)
von: Heng, Liang, et al.
Veröffentlicht: (2025)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
PINNsAgent: Automated PDE Surrogation with Large Language Models
von: Wuwu, Qingpo, et al.
Veröffentlicht: (2025)
von: Wuwu, Qingpo, et al.
Veröffentlicht: (2025)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
von: Li, Haozhan, et al.
Veröffentlicht: (2025)
von: Li, Haozhan, et al.
Veröffentlicht: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
von: Wang, Fangyuan, et al.
Veröffentlicht: (2026)
von: Wang, Fangyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025) -
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2025) -
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2026) -
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025) -
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)