Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Palenicek, Daniel, Vogt, Florian, Watson, Joe, Peters, Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling CrossQ with Weight Normalization
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025)
Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
von: Scherer, Christian, et al.
Veröffentlicht: (2026)
von: Scherer, Christian, et al.
Veröffentlicht: (2026)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies
von: Palenicek, Daniel, et al.
Veröffentlicht: (2026)
von: Palenicek, Daniel, et al.
Veröffentlicht: (2026)
Scalable On-Policy Reinforcement Learning via Adaptive Batch Scaling
von: Park, Jongchan
Veröffentlicht: (2026)
von: Park, Jongchan
Veröffentlicht: (2026)
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
von: Bhatt, Aditya, et al.
Veröffentlicht: (2019)
von: Bhatt, Aditya, et al.
Veröffentlicht: (2019)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
Massively Scaling Explicit Policy-conditioned Value Functions
von: Bohlinger, Nico, et al.
Veröffentlicht: (2025)
von: Bohlinger, Nico, et al.
Veröffentlicht: (2025)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
StableGrad: Backward Scale Control without Batch Normalization
von: Mestre, Jose I., et al.
Veröffentlicht: (2026)
von: Mestre, Jose I., et al.
Veröffentlicht: (2026)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning
von: Zhuang, Yuan, et al.
Veröffentlicht: (2026)
von: Zhuang, Yuan, et al.
Veröffentlicht: (2026)
Making Batch Normalization Great in Federated Deep Learning
von: Zhong, Jike, et al.
Veröffentlicht: (2023)
von: Zhong, Jike, et al.
Veröffentlicht: (2023)
Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search
von: Najib, Amna, et al.
Veröffentlicht: (2024)
von: Najib, Amna, et al.
Veröffentlicht: (2024)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
von: Kim, Donghu, et al.
Veröffentlicht: (2026)
von: Kim, Donghu, et al.
Veröffentlicht: (2026)
Riemannian Batch Normalization: A Gyro Approach
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
Batch Normalization Amplifies Memorization and Privacy Risks
von: Doan, Ngoc Phu, et al.
Veröffentlicht: (2026)
von: Doan, Ngoc Phu, et al.
Veröffentlicht: (2026)
Search-Based Adversarial Estimates for Improving Sample Efficiency in Off-Policy Reinforcement Learning
von: Malato, Federico, et al.
Veröffentlicht: (2025)
von: Malato, Federico, et al.
Veröffentlicht: (2025)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2023)
Zero-Shot Off-Policy Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
Clustering Context in Off-Policy Evaluation
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
Towards Batch-to-Streaming Deep Reinforcement Learning for Continuous Control
von: De Monte, Riccardo, et al.
Veröffentlicht: (2026)
von: De Monte, Riccardo, et al.
Veröffentlicht: (2026)
Learning Action Embeddings for Off-Policy Evaluation
von: Cief, Matej, et al.
Veröffentlicht: (2023)
von: Cief, Matej, et al.
Veröffentlicht: (2023)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
Off-Policy Correction For Multi-Agent Reinforcement Learning
von: Zawalski, Michał, et al.
Veröffentlicht: (2021)
von: Zawalski, Michał, et al.
Veröffentlicht: (2021)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Policy Learning for Off-Dynamics RL with Deficient Support
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
von: Van, Linh Le Pham, et al.
Veröffentlicht: (2024)
Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
von: Kekić, Armin, et al.
Veröffentlicht: (2025)
von: Kekić, Armin, et al.
Veröffentlicht: (2025)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Value-Distributional Model-Based Reinforcement Learning
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
von: Fei, Wen, et al.
Veröffentlicht: (2020)
von: Fei, Wen, et al.
Veröffentlicht: (2020)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
von: Shitanda, Naoki, et al.
Veröffentlicht: (2026)
von: Shitanda, Naoki, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling CrossQ with Weight Normalization
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025) -
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
von: Palenicek, Daniel, et al.
Veröffentlicht: (2025) -
Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
von: Scherer, Christian, et al.
Veröffentlicht: (2026) -
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2024) -
XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies
von: Palenicek, Daniel, et al.
Veröffentlicht: (2026)