ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Weisheng, Lan, Kai, Zhou, Jianyi, Zhao, Bo, Su, Xiu, Tong, Junwen, Guan, Weili, Yang, Shuo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inject Once Survive Later: Backdooring Vision-Language-Action Models to Persist Through Downstream Fine-tuning
by: Zhou, Jianyi, et al.
Published: (2026)
by: Zhou, Jianyi, et al.
Published: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
by: Xiao, Junjin, et al.
Published: (2026)
by: Xiao, Junjin, et al.
Published: (2026)
CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
by: Lee, Sung-Wook, et al.
Published: (2025)
by: Lee, Sung-Wook, et al.
Published: (2025)
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
by: Wang, Xiaofan, et al.
Published: (2025)
by: Wang, Xiaofan, et al.
Published: (2025)
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
by: Xie, Sicheng, et al.
Published: (2025)
by: Xie, Sicheng, et al.
Published: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
by: Li, Yajie, et al.
Published: (2026)
by: Li, Yajie, et al.
Published: (2026)
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
by: Ma, Wenxuan, et al.
Published: (2026)
by: Ma, Wenxuan, et al.
Published: (2026)
Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction
by: Jiang, Shuo, et al.
Published: (2025)
by: Jiang, Shuo, et al.
Published: (2025)
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Latent Action Diffusion for Cross-Embodiment Manipulation
by: Bauer, Erik, et al.
Published: (2025)
by: Bauer, Erik, et al.
Published: (2025)
Contrast, Imitate, Adapt: Learning Robotic Skills From Raw Human Videos
by: Qian, Zhifeng, et al.
Published: (2024)
by: Qian, Zhifeng, et al.
Published: (2024)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
by: Zhou, Pengfei, et al.
Published: (2026)
by: Zhou, Pengfei, et al.
Published: (2026)
TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video
by: Zhou, Jianyi, et al.
Published: (2026)
by: Zhou, Jianyi, et al.
Published: (2026)
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
by: Songwei, Wu, et al.
Published: (2026)
by: Songwei, Wu, et al.
Published: (2026)
DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
by: Yang, Zebin, et al.
Published: (2026)
by: Yang, Zebin, et al.
Published: (2026)
Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation
by: Yao, Xiangtong, et al.
Published: (2023)
by: Yao, Xiangtong, et al.
Published: (2023)
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
by: Tian, Yuxuan, et al.
Published: (2026)
by: Tian, Yuxuan, et al.
Published: (2026)
ACORN: Adaptive Contrastive Optimization for Safe and Robust Fine-Grained Robotic Manipulation
by: Zhou, Zhongquan, et al.
Published: (2025)
by: Zhou, Zhongquan, et al.
Published: (2025)
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
Instruct2Act: From Human Instruction to Actions Sequencing and Execution via Robot Action Network for Robotic Manipulation
by: Sharma, Archit, et al.
Published: (2026)
by: Sharma, Archit, et al.
Published: (2026)
ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue
by: Zhang, Daoxuan, et al.
Published: (2026)
by: Zhang, Daoxuan, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
by: Bi, Hongzhe, et al.
Published: (2025)
by: Bi, Hongzhe, et al.
Published: (2025)
Physically Consistent Humanoid Loco-Manipulation using Latent Diffusion Models
by: Taouil, Ilyass, et al.
Published: (2025)
by: Taouil, Ilyass, et al.
Published: (2025)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
by: Tan, Huajie, et al.
Published: (2026)
by: Tan, Huajie, et al.
Published: (2026)
UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations
by: Yuan, Tingyu, et al.
Published: (2025)
by: Yuan, Tingyu, et al.
Published: (2025)
Language-Grounded Decoupled Action Representation for Robotic Manipulation
by: Weng, Wuding, et al.
Published: (2026)
by: Weng, Wuding, et al.
Published: (2026)
Generalist Robot Manipulation beyond Action Labeled Data
by: Spiridonov, Alexander, et al.
Published: (2025)
by: Spiridonov, Alexander, et al.
Published: (2025)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
by: Ma, Teli, et al.
Published: (2026)
by: Ma, Teli, et al.
Published: (2026)
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation
by: Niu, Ye, et al.
Published: (2025)
by: Niu, Ye, et al.
Published: (2025)
MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
by: Shah, Rutav, et al.
Published: (2025)
by: Shah, Rutav, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
ACDC: Adaptive Curriculum Planning with Dynamic Contrastive Control for Goal-Conditioned Reinforcement Learning in Robotic Manipulation
by: Wang, Xuerui, et al.
Published: (2026)
by: Wang, Xuerui, et al.
Published: (2026)
EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models
by: Song, Mingchen, et al.
Published: (2026)
by: Song, Mingchen, et al.
Published: (2026)
Haptic Communication in Human-Human and Human-Robot Co-Manipulation
by: Allen, Katherine H., et al.
Published: (2025)
by: Allen, Katherine H., et al.
Published: (2025)
Multimodal Spiking Neural Network for Space Robotic Manipulation
by: Zhang, Liwen, et al.
Published: (2025)
by: Zhang, Liwen, et al.
Published: (2025)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Similar Items
-
Inject Once Survive Later: Backdooring Vision-Language-Action Models to Persist Through Downstream Fine-tuning
by: Zhou, Jianyi, et al.
Published: (2026) -
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026) -
Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
by: Xiao, Junjin, et al.
Published: (2026) -
CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
by: Lee, Sung-Wook, et al.
Published: (2025) -
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
by: Wang, Xiaofan, et al.
Published: (2025)