Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ke, Yijing, Zhang, Zihan, Wang, Ruosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2026)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2026)
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
von: Liu, Junyan, et al.
Veröffentlicht: (2024)
von: Liu, Junyan, et al.
Veröffentlicht: (2024)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
Q-WSL: Optimizing Goal-Conditioned RL with Weighted Supervised Learning via Dynamic Programming
von: Lei, Xing, et al.
Veröffentlicht: (2024)
von: Lei, Xing, et al.
Veröffentlicht: (2024)
Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning
von: Vu, Minh, et al.
Veröffentlicht: (2025)
von: Vu, Minh, et al.
Veröffentlicht: (2025)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
von: Wang, Shengbo
Veröffentlicht: (2026)
von: Wang, Shengbo
Veröffentlicht: (2026)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
Robust Regularized Policy Iteration under Transition Uncertainty
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2025)
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2025)
RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
von: Arzhantsev, Aleksei, et al.
Veröffentlicht: (2025)
von: Arzhantsev, Aleksei, et al.
Veröffentlicht: (2025)
Generating Physical Dynamics under Priors
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
von: Zhou, Zihan, et al.
Veröffentlicht: (2024)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
von: Zhang, Xinnan, et al.
Veröffentlicht: (2025)
von: Zhang, Xinnan, et al.
Veröffentlicht: (2025)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
Generalized Fitted Q-Iteration with Clustered Data
von: Hu, Liyuan, et al.
Veröffentlicht: (2025)
von: Hu, Liyuan, et al.
Veröffentlicht: (2025)
Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning
von: Deproost, Senne, et al.
Veröffentlicht: (2025)
von: Deproost, Senne, et al.
Veröffentlicht: (2025)
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2023)
von: Ding, Zihan, et al.
Veröffentlicht: (2023)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL
von: Lei, Xing, et al.
Veröffentlicht: (2026)
von: Lei, Xing, et al.
Veröffentlicht: (2026)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning
von: Manda, Kausthubh, et al.
Veröffentlicht: (2025)
von: Manda, Kausthubh, et al.
Veröffentlicht: (2025)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
Fitted Q-Iteration via Max-Plus-Linear Approximation
von: Liu, Y., et al.
Veröffentlicht: (2024)
von: Liu, Y., et al.
Veröffentlicht: (2024)
Deterministic Bounds and Random Estimates of Metric Tensors on Neuromanifolds
von: Sun, Ke
Veröffentlicht: (2025)
von: Sun, Ke
Veröffentlicht: (2025)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
von: Zhao, Guangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Guangyu, et al.
Veröffentlicht: (2024)
Policy Learning for Off-Dynamics RL with Deficient Support
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
von: He, Longxiang, et al.
Veröffentlicht: (2025)
von: He, Longxiang, et al.
Veröffentlicht: (2025)
Recursive Backwards Q-Learning in Deterministic Environments
von: Diekhoff, Jan, et al.
Veröffentlicht: (2024)
von: Diekhoff, Jan, et al.
Veröffentlicht: (2024)
Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders
von: Bruns-Smith, David, et al.
Veröffentlicht: (2023)
von: Bruns-Smith, David, et al.
Veröffentlicht: (2023)
Near-Equivalent Q-learning Policies for Dynamic Treatment Regimes
von: Yazzourh, Sophia, et al.
Veröffentlicht: (2026)
von: Yazzourh, Sophia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
von: Wu, Runzhe, et al.
Veröffentlicht: (2024) -
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024) -
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024) -
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2026) -
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
von: Liu, Junyan, et al.
Veröffentlicht: (2024)