StaQ it! Growing neural networks for Policy Mirror Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Shilova, Alena, Davey, Alex, Driss, Brahim, Akrour, Riad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
by: Driss, Brahim, et al.
Published: (2025)
by: Driss, Brahim, et al.
Published: (2025)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
by: Kohler, Hector, et al.
Published: (2025)
by: Kohler, Hector, et al.
Published: (2025)
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
Functional Acceleration for Policy Mirror Descent
by: Chelu, Veronica, et al.
Published: (2024)
by: Chelu, Veronica, et al.
Published: (2024)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
AMStraMGRAM: Adaptive Multi-cutoff Strategy Modification for ANaGRAM
by: Schwencke, Nilo, et al.
Published: (2025)
by: Schwencke, Nilo, et al.
Published: (2025)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
by: Zhang, Jinghan, et al.
Published: (2026)
by: Zhang, Jinghan, et al.
Published: (2026)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
by: Wang, Zeyuan, et al.
Published: (2026)
by: Wang, Zeyuan, et al.
Published: (2026)
Mirror Descent Using the Tempesta Generalized Multi-parametric Logarithms
by: Cichocki, Andrzej
Published: (2025)
by: Cichocki, Andrzej
Published: (2025)
When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning
by: Berthelot, Yann, et al.
Published: (2026)
by: Berthelot, Yann, et al.
Published: (2026)
Adaptive Online Mirror Descent for Tchebycheff Scalarization in Multi-Objective Learning
by: Liu, Meitong, et al.
Published: (2024)
by: Liu, Meitong, et al.
Published: (2024)
Augmented Bayesian Policy Search
by: Kallel, Mahdi, et al.
Published: (2024)
by: Kallel, Mahdi, et al.
Published: (2024)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
Mirror Descent and Novel Exponentiated Gradient Algorithms Using Trace-Form Entropies and Deformed Logarithms
by: Cichocki, Andrzej, et al.
Published: (2025)
by: Cichocki, Andrzej, et al.
Published: (2025)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training
by: Zhang, Johnny R., et al.
Published: (2025)
by: Zhang, Johnny R., et al.
Published: (2025)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Generalized Euler Logarithm and its Applications in Machine Learning: Natural Gradient, Backpropagation, Generalized EG, Mirror Descent and OLPS
by: Cichocki, Andrzej
Published: (2025)
by: Cichocki, Andrzej
Published: (2025)
StaTS: Spectral Trajectory Schedule Learning for Adaptive Time Series Forecasting with Frequency Guided Denoiser
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution
by: Seyde, Tim, et al.
Published: (2024)
by: Seyde, Tim, et al.
Published: (2024)
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
by: Xu, Hang, et al.
Published: (2024)
by: Xu, Hang, et al.
Published: (2024)
Sobolev acceleration for neural networks
by: Oh, Jong Kwon, et al.
Published: (2025)
by: Oh, Jong Kwon, et al.
Published: (2025)
On permutation-invariant neural networks
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
Attention mechanisms in neural networks
by: Hays, Hasi
Published: (2026)
by: Hays, Hasi
Published: (2026)
Linearity-based neural network compression
by: Dobler, Silas, et al.
Published: (2025)
by: Dobler, Silas, et al.
Published: (2025)
Principles of Lipschitz continuity in neural networks
by: Luo, Róisín
Published: (2026)
by: Luo, Róisín
Published: (2026)
Applying graph neural network to SupplyGraph for supply chain network
by: Han, Kihwan
Published: (2024)
by: Han, Kihwan
Published: (2024)
Understanding the dynamics of the frequency bias in neural networks
by: Molina, Juan, et al.
Published: (2024)
by: Molina, Juan, et al.
Published: (2024)
Graph neural networks informed locally by thermodynamics
by: Tierz, Alicia, et al.
Published: (2024)
by: Tierz, Alicia, et al.
Published: (2024)
Graph neural networks and non-commuting operators
by: Velasco, Mauricio, et al.
Published: (2024)
by: Velasco, Mauricio, et al.
Published: (2024)
LayerCollapse: Adaptive compression of neural networks
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
Configurable Mirror Descent: Towards a Unification of Decision Making
by: Li, Pengdeng, et al.
Published: (2024)
by: Li, Pengdeng, et al.
Published: (2024)
Plasticity as the Mirror of Empowerment
by: Abel, David, et al.
Published: (2025)
by: Abel, David, et al.
Published: (2025)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Similar Items
-
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
by: Driss, Brahim, et al.
Published: (2025) -
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
by: Kohler, Hector, et al.
Published: (2023) -
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
by: Kohler, Hector, et al.
Published: (2025) -
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024) -
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)