Reward-Directed Score-Based Diffusion Models via q-Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Xuefeng, Zha, Jiale, Zhou, Xun Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2025)
by: Huang, Yilie, et al.
Published: (2025)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Sublinear Regret for a Class of Continuous-Time Linear-Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2024)
by: Huang, Yilie, et al.
Published: (2024)
Reinforcement Learning for Jump-Diffusions, with Financial Applications
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Logarithmic regret bounds for continuous-time average-reward Markov decision processes
by: Gao, Xuefeng, et al.
Published: (2022)
by: Gao, Xuefeng, et al.
Published: (2022)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Optimal and Diffusion Transports in Machine Learning
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025)
by: Sheng, Jiayuan, et al.
Published: (2025)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
by: Li, Hongpei, et al.
Published: (2024)
by: Li, Hongpei, et al.
Published: (2024)
Geometric Data Valuation via Leverage Scores
by: Mendoza-Smith, Rodrigo
Published: (2025)
by: Mendoza-Smith, Rodrigo
Published: (2025)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
by: Li, Gang, et al.
Published: (2024)
by: Li, Gang, et al.
Published: (2024)
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
by: Skifstad, Julian, et al.
Published: (2026)
by: Skifstad, Julian, et al.
Published: (2026)
Regret of exploratory policy improvement and $q$-learning
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Hyperparameter Optimization for Driving Strategies Based on Reinforcement Learning
by: Adde, Nihal Acharya, et al.
Published: (2024)
by: Adde, Nihal Acharya, et al.
Published: (2024)
TaskMet: Task-Driven Metric Learning for Model Learning
by: Bansal, Dishank, et al.
Published: (2023)
by: Bansal, Dishank, et al.
Published: (2023)
Convex and Bilevel Optimization for Neuro-Symbolic Inference and Learning
by: Dickens, Charles, et al.
Published: (2024)
by: Dickens, Charles, et al.
Published: (2024)
Accelerating Cutting-Plane Algorithms via Reinforcement Learning Surrogates
by: Mana, Kyle, et al.
Published: (2023)
by: Mana, Kyle, et al.
Published: (2023)
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Joint Problems in Learning Multiple Dynamical Systems
by: Niu, Mengjia, et al.
Published: (2023)
by: Niu, Mengjia, et al.
Published: (2023)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
by: Cui, Mingxuan, et al.
Published: (2025)
by: Cui, Mingxuan, et al.
Published: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Closing the Loop: Coordinating Inventory and Recommendation via Deep Reinforcement Learning on Multiple Timescales
by: Jiang, Jinyang, et al.
Published: (2025)
by: Jiang, Jinyang, et al.
Published: (2025)
Learning the Riccati solution operator for time-varying LQR via Deep Operator Networks
by: Chen, Jun, et al.
Published: (2026)
by: Chen, Jun, et al.
Published: (2026)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Learning Concave Bid Shading Strategies in Online Auctions via Measure-valued Proximal Optimization
by: Nodozi, Iman, et al.
Published: (2025)
by: Nodozi, Iman, et al.
Published: (2025)
To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO
by: Qiu, Zi-Hao, et al.
Published: (2024)
by: Qiu, Zi-Hao, et al.
Published: (2024)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Similar Items
-
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024) -
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2025) -
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025) -
Sublinear Regret for a Class of Continuous-Time Linear-Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2024) -
Reinforcement Learning for Jump-Diffusions, with Financial Applications
by: Gao, Xuefeng, et al.
Published: (2024)