Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Scherer, Christian, Watson, Joe, Gruner, Theo, Palenicek, Daniel, Posner, Ingmar, Peters, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies
by: Palenicek, Daniel, et al.
Published: (2026)
by: Palenicek, Daniel, et al.
Published: (2026)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
by: Flynn, Hamish, et al.
Published: (2026)
by: Flynn, Hamish, et al.
Published: (2026)
Disentangling Dynamical Systems: Causal Representation Learning Meets Local Sparse Attention
by: Baumgartner, Markus W., et al.
Published: (2026)
by: Baumgartner, Markus W., et al.
Published: (2026)
Reward-Free Curricula for Training Robust World Models
by: Rigter, Marc, et al.
Published: (2023)
by: Rigter, Marc, et al.
Published: (2023)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Scaling CrossQ with Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
World Models via Policy-Guided Trajectory Diffusion
by: Rigter, Marc, et al.
Published: (2023)
by: Rigter, Marc, et al.
Published: (2023)
Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion
by: Bohlinger, Nico, et al.
Published: (2025)
by: Bohlinger, Nico, et al.
Published: (2025)
Towards Safe Robot Foundation Models
by: Tölle, Maximilian, et al.
Published: (2025)
by: Tölle, Maximilian, et al.
Published: (2025)
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
by: Kim, Donghu, et al.
Published: (2026)
by: Kim, Donghu, et al.
Published: (2026)
Analysing the Interplay of Vision and Touch for Dexterous Insertion Tasks
by: Lenz, Janis, et al.
Published: (2024)
by: Lenz, Janis, et al.
Published: (2024)
Compete and Compose: Learning Independent Mechanisms for Modular World Models
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Diminishing Return of Value Expansion Methods
by: Palenicek, Daniel, et al.
Published: (2024)
by: Palenicek, Daniel, et al.
Published: (2024)
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
by: Bhatt, Aditya, et al.
Published: (2019)
by: Bhatt, Aditya, et al.
Published: (2019)
Towards Safe Robot Foundation Models Using Inductive Biases
by: Tölle, Maximilian, et al.
Published: (2025)
by: Tölle, Maximilian, et al.
Published: (2025)
DIME:Diffusion-Based Maximum Entropy Reinforcement Learning
by: Celik, Onur, et al.
Published: (2025)
by: Celik, Onur, et al.
Published: (2025)
DDO-RM: Distribution-Level Policy Improvement after Reward Learning
by: Zhang, Tiantian, et al.
Published: (2026)
by: Zhang, Tiantian, et al.
Published: (2026)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
COMBO-Grasp: Learning Constraint-Based Manipulation for Bimanual Occluded Grasping
by: Yamada, Jun, et al.
Published: (2025)
by: Yamada, Jun, et al.
Published: (2025)
Off-Policy Evaluation for Recommendations with Missing-Not-At-Random Rewards
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
by: Ackermann, Johannes, et al.
Published: (2025)
by: Ackermann, Johannes, et al.
Published: (2025)
Efficient Stochastic Optimal Control through Approximate Bayesian Input Inference
by: Watson, Joe, et al.
Published: (2021)
by: Watson, Joe, et al.
Published: (2021)
D-Cubed: Latent Diffusion Trajectory Optimisation for Dexterous Deformable Manipulation
by: Yamada, Jun, et al.
Published: (2024)
by: Yamada, Jun, et al.
Published: (2024)
Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models
by: Peysakhovich, Alexander, et al.
Published: (2026)
by: Peysakhovich, Alexander, et al.
Published: (2026)
Machine Learning with Physics Knowledge for Prediction: A Survey
by: Watson, Joe, et al.
Published: (2024)
by: Watson, Joe, et al.
Published: (2024)
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
Gaitor: Learning a Unified Representation Across Gaits for Real-World Quadruped Locomotion
by: Mitchell, Alexander L., et al.
Published: (2024)
by: Mitchell, Alexander L., et al.
Published: (2024)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling
by: Nikulkov, Alex
Published: (2026)
by: Nikulkov, Alex
Published: (2026)
Off-Policy Value-Based Reinforcement Learning for Large Language Models
by: Wang, Peng-Yuan, et al.
Published: (2026)
by: Wang, Peng-Yuan, et al.
Published: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
Learning Tactile Insertion in the Real World
by: Palenicek, Daniel, et al.
Published: (2024)
by: Palenicek, Daniel, et al.
Published: (2024)
Off-Policy Reinforcement Learning with High Dimensional Reward
by: Lee, Dong Neuck, et al.
Published: (2024)
by: Lee, Dong Neuck, et al.
Published: (2024)
Policy Improvement Reinforcement Learning
by: Wang, Huaiyang, et al.
Published: (2026)
by: Wang, Huaiyang, et al.
Published: (2026)
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
by: Luu, Tung M., et al.
Published: (2025)
by: Luu, Tung M., et al.
Published: (2025)
The Complexity Dynamics of Grokking
by: DeMoss, Branton, et al.
Published: (2024)
by: DeMoss, Branton, et al.
Published: (2024)
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Similar Items
-
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
by: Palenicek, Daniel, et al.
Published: (2025) -
XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies
by: Palenicek, Daniel, et al.
Published: (2026) -
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025) -
Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
by: Flynn, Hamish, et al.
Published: (2026) -
Disentangling Dynamical Systems: Causal Representation Learning Meets Local Sparse Attention
by: Baumgartner, Markus W., et al.
Published: (2026)