LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhuoyang, Liu, Jiaming, Chen, Hao, Yu, Jiale, Guo, Ziyu, Hou, Chengkai, Gu, Chenyang, Mi, Xiangju, Zhang, Renrui, Wu, Kun, Che, Zhengping, Tang, Jian, Heng, Pheng-Ann, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation
von: He, Jingyang, et al.
Veröffentlicht: (2026)
von: He, Jingyang, et al.
Veröffentlicht: (2026)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
von: Guo, Ziyu, et al.
Veröffentlicht: (2024)
von: Guo, Ziyu, et al.
Veröffentlicht: (2024)
Point Cloud Understanding via Attention-Driven Contrastive Learning
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
von: Wei, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wei, Xiaobao, et al.
Veröffentlicht: (2023)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence
von: Hou, Chengkai, et al.
Veröffentlicht: (2025)
von: Hou, Chengkai, et al.
Veröffentlicht: (2025)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
Capabilities and Fundamental Limits of Latent Chain-of-Thought
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
Uncovering Latent Chain of Thought Vectors in Language Models
von: Zhang, Jason, et al.
Veröffentlicht: (2024)
von: Zhang, Jason, et al.
Veröffentlicht: (2024)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
von: Liu, Dong, et al.
Veröffentlicht: (2026)
von: Liu, Dong, et al.
Veröffentlicht: (2026)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026) -
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
von: Gu, Chenyang, et al.
Veröffentlicht: (2025) -
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025) -
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
von: Luo, Yuechen, et al.
Veröffentlicht: (2026) -
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)