Aligning Text-to-Image Diffusion Models with Reward Backpropagation
Fuente:
arXiv
Saved in:
| Main Authors: | Prabhudesai, Mihir, Goyal, Anirudh, Pathak, Deepak, Fragkiadaki, Katerina |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)
by: Prabhudesai, Mihir, et al.
Published: (2024)
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025)
by: Swerdlow, Alexander, et al.
Published: (2025)
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
by: Prabhudesai, Mihir, et al.
Published: (2026)
by: Prabhudesai, Mihir, et al.
Published: (2026)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
by: Ke, Tsung-Wei, et al.
Published: (2024)
by: Ke, Tsung-Wei, et al.
Published: (2024)
Self-Questioning Language Models
by: Chen, Lili, et al.
Published: (2025)
by: Chen, Lili, et al.
Published: (2025)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
ODIN: A Single Model for 2D and 3D Segmentation
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
by: Wang, Yufei, et al.
Published: (2023)
by: Wang, Yufei, et al.
Published: (2023)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
by: Goyal, Divyanshu, et al.
Published: (2026)
by: Goyal, Divyanshu, et al.
Published: (2026)
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
by: Jangir, Yash, et al.
Published: (2025)
by: Jangir, Yash, et al.
Published: (2025)
Neural MP: A Generalist Neural Motion Planner
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
SAPG: Split and Aggregate Policy Gradients
by: Singla, Jayesh, et al.
Published: (2024)
by: Singla, Jayesh, et al.
Published: (2024)
IFG: Internet-Scale Guidance for Functional Grasping Generation
by: Liu, Ray Muxin, et al.
Published: (2025)
by: Liu, Ray Muxin, et al.
Published: (2025)
FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learning
by: Liu, Jason Jingzhou, et al.
Published: (2025)
by: Liu, Jason Jingzhou, et al.
Published: (2025)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
by: Gupta, Gunshi, et al.
Published: (2024)
by: Gupta, Gunshi, et al.
Published: (2024)
Maximizing Confidence Alone Improves Reasoning
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
by: Kuang, Yuxuan, et al.
Published: (2026)
by: Kuang, Yuxuan, et al.
Published: (2026)
ViPRA: Video Prediction for Robot Actions
by: Routray, Sandeep, et al.
Published: (2025)
by: Routray, Sandeep, et al.
Published: (2025)
Adaptive Mobile Manipulation for Articulated Objects In the Open World
by: Xiong, Haoyu, et al.
Published: (2024)
by: Xiong, Haoyu, et al.
Published: (2024)
Bimanual Dexterity for Complex Tasks
by: Shaw, Kenneth, et al.
Published: (2024)
by: Shaw, Kenneth, et al.
Published: (2024)
Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards
by: Yu, Xuehui, et al.
Published: (2026)
by: Yu, Xuehui, et al.
Published: (2026)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
Continuously Improving Mobile Manipulation with Autonomous Real-World RL
by: Mendonca, Russell, et al.
Published: (2024)
by: Mendonca, Russell, et al.
Published: (2024)
SPIN: Simultaneous Perception, Interaction and Navigation
by: Uppal, Shagun, et al.
Published: (2024)
by: Uppal, Shagun, et al.
Published: (2024)
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation
by: Zhang, Zihui, et al.
Published: (2026)
by: Zhang, Zihui, et al.
Published: (2026)
Fractional Diffusion Bridge Models
by: Nobis, Gabriel, et al.
Published: (2025)
by: Nobis, Gabriel, et al.
Published: (2025)
DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
by: Tao, Tony, et al.
Published: (2025)
by: Tao, Tony, et al.
Published: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
Deep Reactive Policy: Learning Reactive Manipulator Motion Planning for Dynamic Environments
by: Yang, Jiahui, et al.
Published: (2025)
by: Yang, Jiahui, et al.
Published: (2025)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
by: Huang, Renming, et al.
Published: (2024)
by: Huang, Renming, et al.
Published: (2024)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
by: Liu, Songming, et al.
Published: (2024)
by: Liu, Songming, et al.
Published: (2024)
Text-to-CAD Evaluation with CADTests
by: Mallis, Dimitrios, et al.
Published: (2026)
by: Mallis, Dimitrios, et al.
Published: (2026)
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
by: Sarch, Gabriel, et al.
Published: (2024)
by: Sarch, Gabriel, et al.
Published: (2024)
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
by: Albaba, Mert, et al.
Published: (2025)
by: Albaba, Mert, et al.
Published: (2025)
Similar Items
-
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024) -
Unified Multimodal Discrete Diffusion
by: Swerdlow, Alexander, et al.
Published: (2025) -
Diffusion Beats Autoregressive in Data-Constrained Settings
by: Prabhudesai, Mihir, et al.
Published: (2025) -
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026) -
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
by: Prabhudesai, Mihir, et al.
Published: (2026)