One-Step Flow Policy Mirror Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tianyi, Ma, Haitong, Li, Na, Wang, Kai, Dai, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Online Reinforcement Learning for Diffusion Policy
von: Ma, Haitong, et al.
Veröffentlicht: (2025)
von: Ma, Haitong, et al.
Veröffentlicht: (2025)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based Viewpoint
von: Ma, Haitong, et al.
Veröffentlicht: (2024)
von: Ma, Haitong, et al.
Veröffentlicht: (2024)
Efficient Duple Perturbation Robustness in Low-rank MDPs
von: Hu, Yang, et al.
Veröffentlicht: (2024)
von: Hu, Yang, et al.
Veröffentlicht: (2024)
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
von: Gao, Ting, et al.
Veröffentlicht: (2026)
von: Gao, Ting, et al.
Veröffentlicht: (2026)
Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
von: Ma, Haitong, et al.
Veröffentlicht: (2025)
von: Ma, Haitong, et al.
Veröffentlicht: (2025)
On the Effect of Regularization in Policy Mirror Descent
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
von: Kleuker, Jan Felix, et al.
Veröffentlicht: (2025)
Policy Mirror Descent with Lookahead
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
Stochastic Nonlinear Control via Finite-dimensional Spectral Dynamic Embedding
von: Ren, Zhaolin, et al.
Veröffentlicht: (2023)
von: Ren, Zhaolin, et al.
Veröffentlicht: (2023)
On the Convergence of Policy in Unregularized Policy Mirror Descent
von: Lin, Dachao, et al.
Veröffentlicht: (2022)
von: Lin, Dachao, et al.
Veröffentlicht: (2022)
FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies
von: Gao, Chenxiao, et al.
Veröffentlicht: (2026)
von: Gao, Chenxiao, et al.
Veröffentlicht: (2026)
Functional Acceleration for Policy Mirror Descent
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
von: Chelu, Veronica, et al.
Veröffentlicht: (2024)
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds
von: Xu, Zhenghao, et al.
Veröffentlicht: (2023)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2023)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
von: Liu, Jiacai, et al.
Veröffentlicht: (2025)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge
von: Gong, Yuehu, et al.
Veröffentlicht: (2026)
von: Gong, Yuehu, et al.
Veröffentlicht: (2026)
Latent Policy Steering through One-Step Flow Policies
von: Im, Hokyun, et al.
Veröffentlicht: (2026)
von: Im, Hokyun, et al.
Veröffentlicht: (2026)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
von: Zhou, Xubin, et al.
Veröffentlicht: (2026)
von: Zhou, Xubin, et al.
Veröffentlicht: (2026)
Primal-Dual Spectral Representation for Off-policy Evaluation
von: Hu, Yang, et al.
Veröffentlicht: (2024)
von: Hu, Yang, et al.
Veröffentlicht: (2024)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
A Mirror Descent Perspective of Smoothed Sign Descent
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
von: Sherman, Uri, et al.
Veröffentlicht: (2025)
StaQ it! Growing neural networks for Policy Mirror Descent
von: Shilova, Alena, et al.
Veröffentlicht: (2025)
von: Shilova, Alena, et al.
Veröffentlicht: (2025)
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
von: Li, Wenye, et al.
Veröffentlicht: (2025)
von: Li, Wenye, et al.
Veröffentlicht: (2025)
Parameter-free Mirror Descent
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2022)
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2022)
Mirror Descent on Riemannian Manifolds
von: Jiang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Jiang, Jiaxin, et al.
Veröffentlicht: (2026)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
von: Bossens, David M., et al.
Veröffentlicht: (2025)
von: Bossens, David M., et al.
Veröffentlicht: (2025)
Distributed Thompson sampling under constrained communication
von: Zerefa, Saba, et al.
Veröffentlicht: (2024)
von: Zerefa, Saba, et al.
Veröffentlicht: (2024)
The Hidden Cost of Approximation in Online Mirror Descent
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2025)
Trajectory Consistency for One-Step Generation on Euler Mean Flows
von: Li, Zhiqi, et al.
Veröffentlicht: (2026)
von: Li, Zhiqi, et al.
Veröffentlicht: (2026)
Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2026)
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
von: Zhang, Yuyang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2025)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
von: Shu, Yao, et al.
Veröffentlicht: (2026)
von: Shu, Yao, et al.
Veröffentlicht: (2026)
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
von: Zhang, Yilang, et al.
Veröffentlicht: (2025)
von: Zhang, Yilang, et al.
Veröffentlicht: (2025)
Configurable Mirror Descent: Towards a Unification of Decision Making
von: Li, Pengdeng, et al.
Veröffentlicht: (2024)
von: Li, Pengdeng, et al.
Veröffentlicht: (2024)
One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
von: Nguyen, Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh, et al.
Veröffentlicht: (2025)
A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
Value Mirror Descent for Reinforcement Learning
von: Jia, Zhichao, et al.
Veröffentlicht: (2026)
von: Jia, Zhichao, et al.
Veröffentlicht: (2026)
A Mirror Descent-Based Algorithm for Corruption-Tolerant Distributed Gradient Descent
von: Wang, Shuche, et al.
Veröffentlicht: (2024)
von: Wang, Shuche, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Online Reinforcement Learning for Diffusion Policy
von: Ma, Haitong, et al.
Veröffentlicht: (2025) -
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026) -
Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based Viewpoint
von: Ma, Haitong, et al.
Veröffentlicht: (2024) -
Efficient Duple Perturbation Robustness in Low-rank MDPs
von: Hu, Yang, et al.
Veröffentlicht: (2024) -
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
von: Gao, Ting, et al.
Veröffentlicht: (2026)