Saved in:
| Main Authors: | de Vries, Joery A., He, Jinke, Oren, Yaniv, Spaan, Matthijs T. J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.06048 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
by: de Vries, Joery A., et al.
Published: (2025)
by: de Vries, Joery A., et al.
Published: (2025)
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
by: de Vries, Joery A., et al.
Published: (2026)
by: de Vries, Joery A., et al.
Published: (2026)
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
by: Oren, Yaniv, et al.
Published: (2026)
by: Oren, Yaniv, et al.
Published: (2026)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)
by: Murgoci, Vlad, et al.
Published: (2026)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
What model does MuZero learn?
by: He, Jinke, et al.
Published: (2023)
by: He, Jinke, et al.
Published: (2023)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
by: Ribeiro, João G., et al.
Published: (2025)
by: Ribeiro, João G., et al.
Published: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
by: Suau, Miguel, et al.
Published: (2022)
by: Suau, Miguel, et al.
Published: (2022)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
by: Suau, Miguel, et al.
Published: (2023)
by: Suau, Miguel, et al.
Published: (2023)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
by: Mambelli, Davide, et al.
Published: (2024)
by: Mambelli, Davide, et al.
Published: (2024)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
Sparse Masked Attention Policies for Reliable Generalization
by: Horsch, Caroline, et al.
Published: (2026)
by: Horsch, Caroline, et al.
Published: (2026)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Positive Experience Reflection for Agents in Interactive Text Environments
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
by: van der Vaart, Pascal R., et al.
Published: (2025)
by: van der Vaart, Pascal R., et al.
Published: (2025)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Trust Region On-Policy Distillation
by: Xing, Xingrun, et al.
Published: (2026)
by: Xing, Xingrun, et al.
Published: (2026)
Regional Expected Improvement for Efficient Trust Region Selection in High-Dimensional Bayesian Optimization
by: Namura, Nobuo, et al.
Published: (2024)
by: Namura, Nobuo, et al.
Published: (2024)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
by: Trivedi, Gaurish, et al.
Published: (2026)
by: Trivedi, Gaurish, et al.
Published: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
by: Zhang, Shijie, et al.
Published: (2026)
by: Zhang, Shijie, et al.
Published: (2026)
Matrix Low-Rank Trust Region Policy Optimization
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Deep Gaussian Process Proximal Policy Optimization
by: van der Lende, Matthijs, et al.
Published: (2025)
by: van der Lende, Matthijs, et al.
Published: (2025)
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
by: Evers, Thomas, et al.
Published: (2026)
by: Evers, Thomas, et al.
Published: (2026)
QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
by: Lee, Doyeon, et al.
Published: (2026)
by: Lee, Doyeon, et al.
Published: (2026)
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Semi-Supervised Hypothesis Testing by Betting on Predictions
by: Tenzer, Yaniv, et al.
Published: (2026)
by: Tenzer, Yaniv, et al.
Published: (2026)
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
by: Tolochinsky, Elad, et al.
Published: (2026)
by: Tolochinsky, Elad, et al.
Published: (2026)
Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning
by: Wijesundara, Chulabhaya, et al.
Published: (2026)
by: Wijesundara, Chulabhaya, et al.
Published: (2026)
Bayesian Modeling and Estimation of Linear Time-Varying Systems using Neural Networks and Gaussian Processes
by: Shulman, Yaniv
Published: (2025)
by: Shulman, Yaniv
Published: (2025)
Similar Items
-
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
by: de Vries, Joery A., et al.
Published: (2025) -
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
by: de Vries, Joery A., et al.
Published: (2026) -
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
by: Oren, Yaniv, et al.
Published: (2026) -
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025) -
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)