Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
Fuente:
arXiv
Saved in:
| Main Authors: | Vosylius, Vitalis, Seo, Younggyo, Uruç, Jafar, James, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuous Control with Coarse-to-fine Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Instant Policy: In-Context Imitation Learning via Graph Diffusion
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark
by: Chernyadev, Nikita, et al.
Published: (2024)
by: Chernyadev, Nikita, et al.
Published: (2024)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023)
by: Prabhudesai, Mihir, et al.
Published: (2023)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Redundancy-aware Action Spaces for Robot Learning
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation
by: Ma, Xiao, et al.
Published: (2024)
by: Ma, Xiao, et al.
Published: (2024)
Generative Image as Action Models
by: Shridhar, Mohit, et al.
Published: (2024)
by: Shridhar, Mohit, et al.
Published: (2024)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
by: Ke, Tsung-Wei, et al.
Published: (2024)
by: Ke, Tsung-Wei, et al.
Published: (2024)
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
The Ingredients for Robotic Diffusion Transformers
by: Dasari, Sudeep, et al.
Published: (2024)
by: Dasari, Sudeep, et al.
Published: (2024)
Seeking Physics in Diffusion Noise
by: Tang, Chujun, et al.
Published: (2026)
by: Tang, Chujun, et al.
Published: (2026)
Fractional Diffusion Bridge Models
by: Nobis, Gabriel, et al.
Published: (2025)
by: Nobis, Gabriel, et al.
Published: (2025)
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025)
by: Swerdlow, Alexander, et al.
Published: (2025)
Behavioural Cloning in VizDoom
by: Spick, Ryan, et al.
Published: (2024)
by: Spick, Ryan, et al.
Published: (2024)
MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)
by: Prabhudesai, Mihir, et al.
Published: (2024)
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
DiffClone: Enhanced Behaviour Cloning in Robotics with Diffusion-Driven Policy Learning
by: Mani, Sabariswaran, et al.
Published: (2024)
by: Mani, Sabariswaran, et al.
Published: (2024)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
CRAFT: Video Diffusion for Bimanual Robot Data Generation
by: Chen, Jason, et al.
Published: (2026)
by: Chen, Jason, et al.
Published: (2026)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
by: Huang, Renming, et al.
Published: (2024)
by: Huang, Renming, et al.
Published: (2024)
D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
by: Liu, I-Chun Arthur, et al.
Published: (2025)
by: Liu, I-Chun Arthur, et al.
Published: (2025)
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
by: Liu, Songming, et al.
Published: (2024)
by: Liu, Songming, et al.
Published: (2024)
Non-rigid Relative Placement through 3D Dense Diffusion
by: Cai, Eric, et al.
Published: (2024)
by: Cai, Eric, et al.
Published: (2024)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
by: Hu, Zichao, et al.
Published: (2025)
by: Hu, Zichao, et al.
Published: (2025)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
by: Albaba, Mert, et al.
Published: (2025)
by: Albaba, Mert, et al.
Published: (2025)
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
by: He, Wenkun, et al.
Published: (2024)
by: He, Wenkun, et al.
Published: (2024)
DiMSam: Diffusion Models as Samplers for Task and Motion Planning under Partial Observability
by: Fang, Xiaolin, et al.
Published: (2023)
by: Fang, Xiaolin, et al.
Published: (2023)
Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
by: Zhang, Lunjun, et al.
Published: (2023)
by: Zhang, Lunjun, et al.
Published: (2023)
DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
by: Gupta, Gunshi, et al.
Published: (2024)
by: Gupta, Gunshi, et al.
Published: (2024)
TempBEV: Improving Learned BEV Encoders with Combined Image and BEV Space Temporal Aggregation
by: Monninger, Thomas, et al.
Published: (2024)
by: Monninger, Thomas, et al.
Published: (2024)
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Green Screen Augmentation Enables Scene Generalisation in Robotic Manipulation
by: Teoh, Eugene, et al.
Published: (2024)
by: Teoh, Eugene, et al.
Published: (2024)
Similar Items
-
Continuous Control with Coarse-to-fine Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024) -
Instant Policy: In-Context Imitation Learning via Graph Diffusion
by: Vosylius, Vitalis, et al.
Published: (2024) -
BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark
by: Chernyadev, Nikita, et al.
Published: (2024) -
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023) -
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)