Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Lirui, Chen, Xinlei, Zhao, Jialiang, He, Kaiming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
por: Wang, Lirui, et al.
Publicado: (2025)
por: Wang, Lirui, et al.
Publicado: (2025)
Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
por: Zhao, Jialiang, et al.
Publicado: (2024)
por: Zhao, Jialiang, et al.
Publicado: (2024)
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
por: He, Haoran, et al.
Publicado: (2024)
por: He, Haoran, et al.
Publicado: (2024)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
por: Chen, Xinlei, et al.
Publicado: (2024)
por: Chen, Xinlei, et al.
Publicado: (2024)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
por: Juneja, Shubham, et al.
Publicado: (2024)
por: Juneja, Shubham, et al.
Publicado: (2024)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
por: Zhang, Haiming, et al.
Publicado: (2024)
por: Zhang, Haiming, et al.
Publicado: (2024)
GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
por: Hua, Pu, et al.
Publicado: (2024)
por: Hua, Pu, et al.
Publicado: (2024)
Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and Online Self-Supervision
por: Mattamala, Matías, et al.
Publicado: (2024)
por: Mattamala, Matías, et al.
Publicado: (2024)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
por: Albaba, Mert, et al.
Publicado: (2025)
por: Albaba, Mert, et al.
Publicado: (2025)
GPD-1: Generative Pre-training for Driving
por: Xie, Zixun, et al.
Publicado: (2024)
por: Xie, Zixun, et al.
Publicado: (2024)
Multi-Camera View Scaling for Data-Efficient Robot Imitation Learning
por: Xie, Yichen, et al.
Publicado: (2026)
por: Xie, Yichen, et al.
Publicado: (2026)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
por: Wang, Lirui, et al.
Publicado: (2023)
por: Wang, Lirui, et al.
Publicado: (2023)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
por: Zhang, Junjie, et al.
Publicado: (2024)
por: Zhang, Junjie, et al.
Publicado: (2024)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
por: Wang, Junlin, et al.
Publicado: (2025)
por: Wang, Junlin, et al.
Publicado: (2025)
Visual Whole-Body Control for Legged Loco-Manipulation
por: Liu, Minghuan, et al.
Publicado: (2024)
por: Liu, Minghuan, et al.
Publicado: (2024)
Transformers without Normalization
por: Zhu, Jiachen, et al.
Publicado: (2025)
por: Zhu, Jiachen, et al.
Publicado: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
por: Wang, Fangyuan, et al.
Publicado: (2026)
por: Wang, Fangyuan, et al.
Publicado: (2026)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
por: Tsagkas, Nikolaos, et al.
Publicado: (2025)
por: Tsagkas, Nikolaos, et al.
Publicado: (2025)
Latent Representations for Visual Proprioception in Inexpensive Robots
por: Sheikholeslami, Sahara, et al.
Publicado: (2025)
por: Sheikholeslami, Sahara, et al.
Publicado: (2025)
TraIL-Det: Transformation-Invariant Local Feature Networks for 3D LiDAR Object Detection with Unsupervised Pre-Training
por: Li, Li, et al.
Publicado: (2024)
por: Li, Li, et al.
Publicado: (2024)
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model
por: Endres, Jannik, et al.
Publicado: (2025)
por: Endres, Jannik, et al.
Publicado: (2025)
Light-SLAM: A Robust Deep-Learning Visual SLAM System Based on LightGlue under Challenging Lighting Conditions
por: Zhao, Zhiqi, et al.
Publicado: (2024)
por: Zhao, Zhiqi, et al.
Publicado: (2024)
DOGE: An Extrinsic Orientation and Gyroscope Bias Estimation for Visual-Inertial Odometry Initialization
por: Xu, Zewen, et al.
Publicado: (2024)
por: Xu, Zewen, et al.
Publicado: (2024)
A Recipe for Unbounded Data Augmentation in Visual Reinforcement Learning
por: Almuzairee, Abdulaziz, et al.
Publicado: (2024)
por: Almuzairee, Abdulaziz, et al.
Publicado: (2024)
Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics
por: Almuzairee, Abdulaziz, et al.
Publicado: (2026)
por: Almuzairee, Abdulaziz, et al.
Publicado: (2026)
Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
por: Almuzairee, Abdulaziz, et al.
Publicado: (2025)
por: Almuzairee, Abdulaziz, et al.
Publicado: (2025)
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
por: Shao, Xingyu, et al.
Publicado: (2026)
por: Shao, Xingyu, et al.
Publicado: (2026)
Knowledge-aware Graph Transformer for Pedestrian Trajectory Prediction
por: Liu, Yu, et al.
Publicado: (2024)
por: Liu, Yu, et al.
Publicado: (2024)
SoloParkour: Constrained Reinforcement Learning for Visual Locomotion from Privileged Experience
por: Chane-Sane, Elliot, et al.
Publicado: (2024)
por: Chane-Sane, Elliot, et al.
Publicado: (2024)
Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
por: Krauss, Henrik, et al.
Publicado: (2025)
por: Krauss, Henrik, et al.
Publicado: (2025)
Exploring Transformer-Augmented LSTM for Temporal and Spatial Feature Learning in Trajectory Prediction
por: Raskoti, Chandra, et al.
Publicado: (2024)
por: Raskoti, Chandra, et al.
Publicado: (2024)
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
por: Hoque, Ryan, et al.
Publicado: (2025)
por: Hoque, Ryan, et al.
Publicado: (2025)
Latent Object Characteristics Recognition with Visual to Haptic-Audio Cross-modal Transfer Learning
por: Saito, Namiko, et al.
Publicado: (2024)
por: Saito, Namiko, et al.
Publicado: (2024)
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
por: Chen, Suzeyu, et al.
Publicado: (2026)
por: Chen, Suzeyu, et al.
Publicado: (2026)
Large Pre-Trained Models for Bimanual Manipulation in 3D
por: Yurchyk, Hanna, et al.
Publicado: (2025)
por: Yurchyk, Hanna, et al.
Publicado: (2025)
Learning Spatial Structure from Pre-Beamforming Per-Antenna Range-Doppler Radar Data via Visibility-Aware Cross-Modal Supervision
por: Sebastian, George, et al.
Publicado: (2026)
por: Sebastian, George, et al.
Publicado: (2026)
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
por: Howell, Owen Lewis, et al.
Publicado: (2025)
por: Howell, Owen Lewis, et al.
Publicado: (2025)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
por: Mishra, Nikhil, et al.
Publicado: (2024)
por: Mishra, Nikhil, et al.
Publicado: (2024)
Efficient Equivariant Transformer for Self-Driving Agent Modeling
por: Xu, Scott, et al.
Publicado: (2026)
por: Xu, Scott, et al.
Publicado: (2026)
Split Adaptation for Pre-trained Vision Transformers
por: Wang, Lixu, et al.
Publicado: (2025)
por: Wang, Lixu, et al.
Publicado: (2025)
Ejemplares similares
-
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
por: Wang, Lirui, et al.
Publicado: (2025) -
Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
por: Zhao, Jialiang, et al.
Publicado: (2024) -
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
por: He, Haoran, et al.
Publicado: (2024) -
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
por: Chen, Xinlei, et al.
Publicado: (2024) -
DINO Pre-training for Vision-based End-to-end Autonomous Driving
por: Juneja, Shubham, et al.
Publicado: (2024)