Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Gunshi, Yadav, Karmesh, Gal, Yarin, Batra, Dhruv, Kira, Zsolt, Lu, Cong, Rudner, Tim G. J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
by: Yadav, Karmesh, et al.
Published: (2025)
by: Yadav, Karmesh, et al.
Published: (2025)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
by: Andrade, Moises, et al.
Published: (2025)
by: Andrade, Moises, et al.
Published: (2025)
HomeRobot: Open-Vocabulary Mobile Manipulation
by: Yenamandra, Sriram, et al.
Published: (2023)
by: Yenamandra, Sriram, et al.
Published: (2023)
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
by: Coholich, Jeremiah, et al.
Published: (2026)
by: Coholich, Jeremiah, et al.
Published: (2026)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device
by: Chhablani, Gunjan, et al.
Published: (2025)
by: Chhablani, Gunjan, et al.
Published: (2025)
ReLIC: A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI
by: Elawady, Ahmad, et al.
Published: (2024)
by: Elawady, Ahmad, et al.
Published: (2024)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
by: Majumdar, Arjun, et al.
Published: (2023)
by: Majumdar, Arjun, et al.
Published: (2023)
Seeing the Unseen: Visual Common Sense for Semantic Placement
by: Ramrakhya, Ram, et al.
Published: (2024)
by: Ramrakhya, Ram, et al.
Published: (2024)
Towards Open-World Mobile Manipulation in Homes: Lessons from the Neurips 2023 HomeRobot Open Vocabulary Mobile Manipulation Challenge
by: Yenamandra, Sriram, et al.
Published: (2024)
by: Yenamandra, Sriram, et al.
Published: (2024)
Barrier Function Overrides For Non-Convex Fixed Wing Flight Control and Self-Driving Cars
by: Squires, Eric, et al.
Published: (2025)
by: Squires, Eric, et al.
Published: (2025)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
Multi-Transmotion: Pre-trained Model for Human Motion Prediction
by: Gao, Yang, et al.
Published: (2024)
by: Gao, Yang, et al.
Published: (2024)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
What do we learn from a large-scale study of pre-trained visual representations in sim and real environments?
by: Silwal, Sneha, et al.
Published: (2023)
by: Silwal, Sneha, et al.
Published: (2023)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
by: Albaba, Mert, et al.
Published: (2025)
by: Albaba, Mert, et al.
Published: (2025)
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024)
by: Xie, Zixun, et al.
Published: (2024)
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
by: Ljungbergh, William, et al.
Published: (2025)
by: Ljungbergh, William, et al.
Published: (2025)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
HRP: Human Affordances for Robotic Pre-Training
by: Srirama, Mohan Kumar, et al.
Published: (2024)
by: Srirama, Mohan Kumar, et al.
Published: (2024)
4D Contrastive Superflows are Dense 3D Representation Learners
by: Xu, Xiang, et al.
Published: (2024)
by: Xu, Xiang, et al.
Published: (2024)
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
by: Khanna, Mukul, et al.
Published: (2024)
by: Khanna, Mukul, et al.
Published: (2024)
MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
by: Li, Bin, et al.
Published: (2026)
by: Li, Bin, et al.
Published: (2026)
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
by: Wang, Lirui, et al.
Published: (2024)
by: Wang, Lirui, et al.
Published: (2024)
Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models
by: Kyaw, Alexander Htet, et al.
Published: (2025)
by: Kyaw, Alexander Htet, et al.
Published: (2025)
LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes
by: Xu, Xiang, et al.
Published: (2025)
by: Xu, Xiang, et al.
Published: (2025)
TřiVis: Versatile, Reliable, and High-Performance Tool for Computing Visibility in Polygonal Environments
by: Mikula, Jan, et al.
Published: (2024)
by: Mikula, Jan, et al.
Published: (2024)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023)
by: Prabhudesai, Mihir, et al.
Published: (2023)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)
by: Juneja, Shubham, et al.
Published: (2024)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
by: Batra, Sumeet, et al.
Published: (2024)
by: Batra, Sumeet, et al.
Published: (2024)
Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object Segmentation
by: Miao, Ziliang, et al.
Published: (2025)
by: Miao, Ziliang, et al.
Published: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
by: Yin, Zhenhan, et al.
Published: (2025)
by: Yin, Zhenhan, et al.
Published: (2025)
VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation
by: Sun, Zhengnan, et al.
Published: (2025)
by: Sun, Zhengnan, et al.
Published: (2025)
N-QR: Natural Quick Response Codes for Multi-Robot Instance Correspondence
by: Glaser, Nathaniel Moore, et al.
Published: (2024)
by: Glaser, Nathaniel Moore, et al.
Published: (2024)
Pre-Trained Masked Image Model for Mobile Robot Navigation
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
Second-order Theory of Mind for Human Teachers and Robot Learners
by: Callaghan, Patrick, et al.
Published: (2025)
by: Callaghan, Patrick, et al.
Published: (2025)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
Similar Items
-
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
by: Yadav, Karmesh, et al.
Published: (2025) -
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025) -
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
by: Andrade, Moises, et al.
Published: (2025) -
HomeRobot: Open-Vocabulary Mobile Manipulation
by: Yenamandra, Sriram, et al.
Published: (2023) -
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
by: Coholich, Jeremiah, et al.
Published: (2026)