Gespeichert in:
| Hauptverfasser: | Gong, Yuehu, Wang, Zeyuan, Chen, Yulin, Ding, Shutong, Zhou, Qingyuan, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.21621 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
Distributional Reinforcement Learning with Diffusion Bridge Critics
von: Ding, Shutong, et al.
Veröffentlicht: (2026)
von: Ding, Shutong, et al.
Veröffentlicht: (2026)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
Variational Online Mirror Descent for Robust Learning in Schrödinger Bridge
von: Han, Dong-Sig, et al.
Veröffentlicht: (2025)
von: Han, Dong-Sig, et al.
Veröffentlicht: (2025)
One-Step Flow Policy Mirror Descent
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
Value Mirror Descent for Reinforcement Learning
von: Jia, Zhichao, et al.
Veröffentlicht: (2026)
von: Jia, Zhichao, et al.
Veröffentlicht: (2026)
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
von: Ding, Shutong, et al.
Veröffentlicht: (2025)
von: Ding, Shutong, et al.
Veröffentlicht: (2025)
Policy Mirror Descent with Lookahead
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
On the Effect of Regularization in Policy Mirror Descent
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
von: Li, Wenye, et al.
Veröffentlicht: (2025)
von: Li, Wenye, et al.
Veröffentlicht: (2025)
On the Convergence of Policy in Unregularized Policy Mirror Descent
von: Lin, Dachao, et al.
Veröffentlicht: (2022)
von: Lin, Dachao, et al.
Veröffentlicht: (2022)
Functional Acceleration for Policy Mirror Descent
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds
von: Xu, Zhenghao, et al.
Veröffentlicht: (2023)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2023)
Mirror Descent on Reproducing Kernel Banach Spaces
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
Mirror and Preconditioned Gradient Descent in Wasserstein Space
von: Bonet, Clément, et al.
Veröffentlicht: (2024)
von: Bonet, Clément, et al.
Veröffentlicht: (2024)
Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
von: Gao, Ting, et al.
Veröffentlicht: (2026)
von: Gao, Ting, et al.
Veröffentlicht: (2026)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
A Mirror Descent Perspective of Smoothed Sign Descent
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
StaQ it! Growing neural networks for Policy Mirror Descent
von: Shilova, Alena, et al.
Veröffentlicht: (2025)
von: Shilova, Alena, et al.
Veröffentlicht: (2025)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
Estimating Individual Dose-Response Curves under Unobserved Confounders from Observational Data
von: Chen, Shutong, et al.
Veröffentlicht: (2024)
von: Chen, Shutong, et al.
Veröffentlicht: (2024)
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
von: Sun, Mingyang, et al.
Veröffentlicht: (2025)
von: Sun, Mingyang, et al.
Veröffentlicht: (2025)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
Parameter-free Mirror Descent
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2022)
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2022)
Mirror Descent on Riemannian Manifolds
von: Jiang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Jiang, Jiaxin, et al.
Veröffentlicht: (2026)
Stress-Aware Learning under KL Drift via Trust-Decayed Mirror Descent
von: Raj, Gabriel Nixon
Veröffentlicht: (2025)
von: Raj, Gabriel Nixon
Veröffentlicht: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
von: Bossens, David M., et al.
Veröffentlicht: (2025)
von: Bossens, David M., et al.
Veröffentlicht: (2025)
Instance Generation for Meta-Black-Box Optimization through Latent Space Reverse Engineering
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Learning Mixtures of Experts with EM: A Mirror Descent Perspective
von: Fruytier, Quentin, et al.
Veröffentlicht: (2024)
von: Fruytier, Quentin, et al.
Veröffentlicht: (2024)
Mirror Descent Actor Critic via Bounded Advantage Learning
von: Iwaki, Ryo
Veröffentlicht: (2025)
von: Iwaki, Ryo
Veröffentlicht: (2025)
Adaptively Perturbed Mirror Descent for Learning in Games
von: Abe, Kenshi, et al.
Veröffentlicht: (2023)
von: Abe, Kenshi, et al.
Veröffentlicht: (2023)
Extreme Value Policy Optimization for Safe Reinforcement Learning
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance
von: Ding, Shutong, et al.
Veröffentlicht: (2026)
von: Ding, Shutong, et al.
Veröffentlicht: (2026)
The Hidden Cost of Approximation in Online Mirror Descent
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026) -
Distributional Reinforcement Learning with Diffusion Bridge Critics
von: Ding, Shutong, et al.
Veröffentlicht: (2026) -
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025) -
Variational Online Mirror Descent for Robust Learning in Schrödinger Bridge
von: Han, Dong-Sig, et al.
Veröffentlicht: (2025) -
One-Step Flow Policy Mirror Descent
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)