4D Visual Pre-training for Robot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Chengkai, Ze, Yanjie, Fu, Yankai, Gao, Zeyu, Hu, Songbo, Yu, Yue, Zhang, Shanghang, Xu, Huazhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023)
by: Huang, Tao, et al.
Published: (2023)
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap Features
by: Hou, Chengkai, et al.
Published: (2024)
by: Hou, Chengkai, et al.
Published: (2024)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization
by: Xu, Guowei, et al.
Published: (2023)
by: Xu, Guowei, et al.
Published: (2023)
Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion
by: Hu, Kaizhe, et al.
Published: (2024)
by: Hu, Kaizhe, et al.
Published: (2024)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
Formula-Supervised Visual-Geometric Pre-training
by: Yamada, Ryosuke, et al.
Published: (2024)
by: Yamada, Ryosuke, et al.
Published: (2024)
VILA: On Pre-training for Visual Language Models
by: Lin, Ji, et al.
Published: (2023)
by: Lin, Ji, et al.
Published: (2023)
PLIP: Language-Image Pre-training for Person Representation Learning
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving
by: Xu, Shaoqing, et al.
Published: (2024)
by: Xu, Shaoqing, et al.
Published: (2024)
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
by: Zhao, Chenhui, et al.
Published: (2025)
by: Zhao, Chenhui, et al.
Published: (2025)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
by: Yuan, Zhecheng, et al.
Published: (2024)
by: Yuan, Zhecheng, et al.
Published: (2024)
DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo
by: Zhu, Junzhe, et al.
Published: (2024)
by: Zhu, Junzhe, et al.
Published: (2024)
X-Capture: An Open-Source Portable Device for Multi-Sensory Learning
by: Clarke, Samuel, et al.
Published: (2025)
by: Clarke, Samuel, et al.
Published: (2025)
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
by: Xue, Le, et al.
Published: (2023)
by: Xue, Le, et al.
Published: (2023)
Boosting Image Restoration via Priors from Pre-trained Models
by: Xu, Xiaogang, et al.
Published: (2024)
by: Xu, Xiaogang, et al.
Published: (2024)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
by: Fu, Yankai, et al.
Published: (2025)
by: Fu, Yankai, et al.
Published: (2025)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
by: Yan, Xiangchao, et al.
Published: (2023)
by: Yan, Xiangchao, et al.
Published: (2023)
DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
by: Min, Chen, et al.
Published: (2024)
by: Min, Chen, et al.
Published: (2024)
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
by: Wang, Lirui, et al.
Published: (2024)
by: Wang, Lirui, et al.
Published: (2024)
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
by: Long, Rujiao, et al.
Published: (2024)
by: Long, Rujiao, et al.
Published: (2024)
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
by: Li, Wenyu, et al.
Published: (2025)
by: Li, Wenyu, et al.
Published: (2025)
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024)
by: Xie, Zixun, et al.
Published: (2024)
Similar Items
-
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023) -
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
by: Ze, Yanjie, et al.
Published: (2024) -
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
by: Wang, Hao, et al.
Published: (2025) -
Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap Features
by: Hou, Chengkai, et al.
Published: (2024) -
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)