Maximize Your Diffusion: A Study into Reward Maximization and Alignment for Diffusion-based Control
Fuente:
arXiv
Saved in:
| Main Authors: | Huh, Dom, Mohapatra, Prasant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-agent Auto-Bidding with Latent Graph Diffusion Models
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
Multi-agent Reinforcement Learning: A Comprehensive Survey
by: Huh, Dom, et al.
Published: (2023)
by: Huh, Dom, et al.
Published: (2023)
Representation Learning For Efficient Deep Multi-Agent Reinforcement Learning
by: Huh, Dom, et al.
Published: (2024)
by: Huh, Dom, et al.
Published: (2024)
Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
Maximally Permissive Reward Machines
by: Varricchione, Giovanni, et al.
Published: (2024)
by: Varricchione, Giovanni, et al.
Published: (2024)
Jackpot! Alignment as a Maximal Lottery
by: Maura-Rivero, Roberto-Rafael, et al.
Published: (2025)
by: Maura-Rivero, Roberto-Rafael, et al.
Published: (2025)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
JurEE not Judges: safeguarding llm interactions with small, specialised Encoder Ensembles
by: Nasrabadi, Dom
Published: (2024)
by: Nasrabadi, Dom
Published: (2024)
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
by: Tejaswi, Atula, et al.
Published: (2026)
by: Tejaswi, Atula, et al.
Published: (2026)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)
by: Prabhudesai, Mihir, et al.
Published: (2024)
Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review
by: Uehara, Masatoshi, et al.
Published: (2025)
by: Uehara, Masatoshi, et al.
Published: (2025)
Convergence Rate Maximization for Split Learning-based Control of EMG Prosthetic Devices
by: Marinova, Matea, et al.
Published: (2024)
by: Marinova, Matea, et al.
Published: (2024)
Maximizing Confidence Alone Improves Reasoning
by: Prabhudesai, Mihir, et al.
Published: (2025)
by: Prabhudesai, Mihir, et al.
Published: (2025)
How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance
by: Huang, Jerry Y., et al.
Published: (2026)
by: Huang, Jerry Y., et al.
Published: (2026)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
by: Jin, Luozhijie, et al.
Published: (2025)
by: Jin, Luozhijie, et al.
Published: (2025)
Diffusion Alignment as Variational Expectation-Maximization
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
by: Hosseintabar, Danial, et al.
Published: (2025)
by: Hosseintabar, Danial, et al.
Published: (2025)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Reward Sharpness-Aware Fine-Tuning for Diffusion Models
by: Kim, Kwanyoung, et al.
Published: (2026)
by: Kim, Kwanyoung, et al.
Published: (2026)
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment
by: Lee, Banseok, et al.
Published: (2026)
by: Lee, Banseok, et al.
Published: (2026)
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
by: Deng, Fei, et al.
Published: (2024)
by: Deng, Fei, et al.
Published: (2024)
Manifold Sampling via Entropy Maximization
by: Braun, Cornelius V., et al.
Published: (2026)
by: Braun, Cornelius V., et al.
Published: (2026)
How to Explore with Belief: State Entropy Maximization in POMDPs
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
GALA: Graph Diffusion-based Alignment with Jigsaw for Source-free Domain Adaptation
by: Luo, Junyu, et al.
Published: (2024)
by: Luo, Junyu, et al.
Published: (2024)
Preference-Based Alignment of Discrete Diffusion Models
by: Borso, Umberto, et al.
Published: (2025)
by: Borso, Umberto, et al.
Published: (2025)
Revisiting Modularity Maximization for Graph Clustering: A Contrastive Learning Perspective
by: Liu, Yunfei, et al.
Published: (2024)
by: Liu, Yunfei, et al.
Published: (2024)
Offline Diversity Maximization Under Imitation Constraints
by: Vlastelica, Marin, et al.
Published: (2023)
by: Vlastelica, Marin, et al.
Published: (2023)
An Expectation-Maximization Algorithm for Domain Adaptation in Gaussian Causal Models
by: Javidian, Mohammad Ali
Published: (2026)
by: Javidian, Mohammad Ali
Published: (2026)
Supervised Distributional Reduction via Optimal Transport and Dependence Maximization
by: Ramesh, Sai-Aakash, et al.
Published: (2026)
by: Ramesh, Sai-Aakash, et al.
Published: (2026)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
by: Kim, Yeongmin, et al.
Published: (2026)
by: Kim, Yeongmin, et al.
Published: (2026)
Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models
by: Park, Youngrok, et al.
Published: (2025)
by: Park, Youngrok, et al.
Published: (2025)
Dynamic Search for Inference-Time Alignment in Diffusion Models
by: Li, Xiner, et al.
Published: (2025)
by: Li, Xiner, et al.
Published: (2025)
Diffusion Model with Representation Alignment for Protein Inverse Folding
by: Wang, Chenglin, et al.
Published: (2024)
by: Wang, Chenglin, et al.
Published: (2024)
Maximize margins for robust splicing detection
by: de Kergunic, Julien Simon, et al.
Published: (2025)
by: de Kergunic, Julien Simon, et al.
Published: (2025)
Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization
by: Modirshanechi, Alireza, et al.
Published: (2026)
by: Modirshanechi, Alireza, et al.
Published: (2026)
Pairwise Calibrated Rewards for Pluralistic Alignment
by: Halpern, Daniel, et al.
Published: (2025)
by: Halpern, Daniel, et al.
Published: (2025)
Similar Items
-
Multi-agent Auto-Bidding with Latent Graph Diffusion Models
by: Huh, Dom, et al.
Published: (2025) -
Multi-agent Reinforcement Learning: A Comprehensive Survey
by: Huh, Dom, et al.
Published: (2023) -
Representation Learning For Efficient Deep Multi-Agent Reinforcement Learning
by: Huh, Dom, et al.
Published: (2024) -
Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
by: Huh, Dom, et al.
Published: (2025) -
Maximally Permissive Reward Machines
by: Varricchione, Giovanni, et al.
Published: (2024)