CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Chubin, Wang, Jianan, Gao, Zifeng, Su, Yue, Dai, Tianru, Zhou, Cai, Lu, Jiwen, Tang, Yansong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
di: Dai, Weisheng, et al.
Pubblicazione: (2026)
di: Dai, Weisheng, et al.
Pubblicazione: (2026)
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
di: Fu, Yiyang, et al.
Pubblicazione: (2026)
di: Fu, Yiyang, et al.
Pubblicazione: (2026)
DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
di: Su, Yue, et al.
Pubblicazione: (2025)
di: Su, Yue, et al.
Pubblicazione: (2025)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
di: Lu, Guanxing, et al.
Pubblicazione: (2025)
Latent Action Pretraining from Videos
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
di: Li, Zuolei, et al.
Pubblicazione: (2025)
di: Li, Zuolei, et al.
Pubblicazione: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
di: Li, Runze, et al.
Pubblicazione: (2026)
di: Li, Runze, et al.
Pubblicazione: (2026)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
di: Luo, Hao, et al.
Pubblicazione: (2025)
di: Luo, Hao, et al.
Pubblicazione: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
di: Li, Qiwei, et al.
Pubblicazione: (2026)
di: Li, Qiwei, et al.
Pubblicazione: (2026)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
di: Tang, Zuojin, et al.
Pubblicazione: (2026)
di: Tang, Zuojin, et al.
Pubblicazione: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
di: Luo, Hao, et al.
Pubblicazione: (2026)
di: Luo, Hao, et al.
Pubblicazione: (2026)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
di: Yang, Rushuai, et al.
Pubblicazione: (2025)
di: Yang, Rushuai, et al.
Pubblicazione: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025)
di: Li, Qixiu, et al.
Pubblicazione: (2025)
Latent Action Pretraining Through World Modeling
di: Tharwat, Bahey, et al.
Pubblicazione: (2025)
di: Tharwat, Bahey, et al.
Pubblicazione: (2025)
Cross-Hand Latent Representation for Vision-Language-Action Models
di: Jiang, Guangqi, et al.
Pubblicazione: (2026)
di: Jiang, Guangqi, et al.
Pubblicazione: (2026)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
di: Govind, Manish Kumar, et al.
Pubblicazione: (2026)
di: Govind, Manish Kumar, et al.
Pubblicazione: (2026)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
di: Liu, Huihan, et al.
Pubblicazione: (2026)
di: Liu, Huihan, et al.
Pubblicazione: (2026)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
di: Chen, Xiaoyu, et al.
Pubblicazione: (2025)
di: Chen, Xiaoyu, et al.
Pubblicazione: (2025)
Toward Embodiment Equivariant Vision-Language-Action Policy
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
di: Grover, Shresth, et al.
Pubblicazione: (2025)
di: Grover, Shresth, et al.
Pubblicazione: (2025)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
di: Guo, Minghao, et al.
Pubblicazione: (2025)
di: Guo, Minghao, et al.
Pubblicazione: (2025)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
di: Dai, Yuntao, et al.
Pubblicazione: (2025)
di: Dai, Yuntao, et al.
Pubblicazione: (2025)
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
di: Liu, Yudong, et al.
Pubblicazione: (2026)
di: Liu, Yudong, et al.
Pubblicazione: (2026)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
di: Nie, Dujun, et al.
Pubblicazione: (2026)
di: Nie, Dujun, et al.
Pubblicazione: (2026)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
di: Chen, Guangyan, et al.
Pubblicazione: (2025)
di: Chen, Guangyan, et al.
Pubblicazione: (2025)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
di: Chen, Guangyan, et al.
Pubblicazione: (2025)
di: Chen, Guangyan, et al.
Pubblicazione: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
di: Lin, Haitao, et al.
Pubblicazione: (2026)
di: Lin, Haitao, et al.
Pubblicazione: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
di: Li, Pengteng, et al.
Pubblicazione: (2026)
di: Li, Pengteng, et al.
Pubblicazione: (2026)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
di: Zhu, Xiang, et al.
Pubblicazione: (2026)
di: Zhu, Xiang, et al.
Pubblicazione: (2026)
Vision in Action: Learning Active Perception from Human Demonstrations
di: Xiong, Haoyu, et al.
Pubblicazione: (2025)
di: Xiong, Haoyu, et al.
Pubblicazione: (2025)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
di: Liu, Zhuoyang, et al.
Pubblicazione: (2026)
di: Liu, Zhuoyang, et al.
Pubblicazione: (2026)
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
di: Lu, Guanxing, et al.
Pubblicazione: (2024)
di: Lu, Guanxing, et al.
Pubblicazione: (2024)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
di: Zhang, Borong, et al.
Pubblicazione: (2025)
di: Zhang, Borong, et al.
Pubblicazione: (2025)
Embodiment Transfer Learning for Vision-Language-Action Models
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
di: Zhang, Hanxin, et al.
Pubblicazione: (2026)
di: Zhang, Hanxin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
di: Dai, Weisheng, et al.
Pubblicazione: (2026) -
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
di: Fu, Yiyang, et al.
Pubblicazione: (2026) -
DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
di: Su, Yue, et al.
Pubblicazione: (2025) -
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
di: Lu, Guanxing, et al.
Pubblicazione: (2025) -
Latent Action Pretraining from Videos
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)