4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiahui, Chen, Yurui, Xu, Yueming, Huang, Ze, Zhou, Yanpeng, Yuan, Yu-Jie, Cai, Xinyue, Huang, Guowei, Quan, Xingyue, Xu, Hang, Zhang, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
von: Xu, Yueming, et al.
Veröffentlicht: (2025)
von: Xu, Yueming, et al.
Veröffentlicht: (2025)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
von: Huang, Helong, et al.
Veröffentlicht: (2025)
von: Huang, Helong, et al.
Veröffentlicht: (2025)
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
von: Huang, Alex S., et al.
Veröffentlicht: (2026)
von: Huang, Alex S., et al.
Veröffentlicht: (2026)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
von: Fu, Yuxia, et al.
Veröffentlicht: (2025)
von: Fu, Yuxia, et al.
Veröffentlicht: (2025)
CRL-VLA: Continual Vision-Language-Action Learning
von: Zeng, Qixin, et al.
Veröffentlicht: (2026)
von: Zeng, Qixin, et al.
Veröffentlicht: (2026)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
von: Luo, Hao, et al.
Veröffentlicht: (2026)
von: Luo, Hao, et al.
Veröffentlicht: (2026)
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Whole-Body Inverse Kinematics with Graph Diffusion
von: Huang, Helong, et al.
Veröffentlicht: (2026)
von: Huang, Helong, et al.
Veröffentlicht: (2026)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
von: Li, Boyu, et al.
Veröffentlicht: (2026)
von: Li, Boyu, et al.
Veröffentlicht: (2026)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
von: Wu, You, et al.
Veröffentlicht: (2026)
von: Wu, You, et al.
Veröffentlicht: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
Spatiotemporal Calibration for Laser Vision Sensor in Hand-eye System Based on Straight-line Constraint
von: Yang, Peiwen, et al.
Veröffentlicht: (2025)
von: Yang, Peiwen, et al.
Veröffentlicht: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
Reinforcing Action Policies by Prophesying
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
von: Han, Gaoge, et al.
Veröffentlicht: (2026)
von: Han, Gaoge, et al.
Veröffentlicht: (2026)
EvoVLA: Self-Evolving Vision-Language-Action Model
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
RynnVLA-002: A Unified Vision-Language-Action and World Model
von: Cen, Jun, et al.
Veröffentlicht: (2025)
von: Cen, Jun, et al.
Veröffentlicht: (2025)
Enhancing Spatiotemporal Resampling with a Novel MIS Weight
von: Xingyue Pan, et al.
Veröffentlicht: (2024)
von: Xingyue Pan, et al.
Veröffentlicht: (2024)
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
von: Zhu, Yijie, et al.
Veröffentlicht: (2026)
von: Zhu, Yijie, et al.
Veröffentlicht: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025) -
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
von: Xu, Yueming, et al.
Veröffentlicht: (2025) -
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
von: Huang, Helong, et al.
Veröffentlicht: (2025) -
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
von: Huang, Alex S., et al.
Veröffentlicht: (2026) -
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
von: Fu, Yuxia, et al.
Veröffentlicht: (2025)