Universal Value-Function Uncertainties
Fuente:
arXiv
Saved in:
| Main Authors: | Zanger, Moritz A., Weltevrede, Max, Oren, Yaniv, Van der Vaart, Pascal R., Horsch, Caroline, Böhmer, Wendelin, Spaan, Matthijs T. J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Sparse Masked Attention Policies for Reliable Generalization
by: Horsch, Caroline, et al.
Published: (2026)
by: Horsch, Caroline, et al.
Published: (2026)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
by: van der Vaart, Pascal R., et al.
Published: (2025)
by: van der Vaart, Pascal R., et al.
Published: (2025)
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
by: Evers, Thomas, et al.
Published: (2026)
by: Evers, Thomas, et al.
Published: (2026)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
by: Ribeiro, João G., et al.
Published: (2025)
by: Ribeiro, João G., et al.
Published: (2025)
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
by: Tamassia, Isidoro, et al.
Published: (2025)
by: Tamassia, Isidoro, et al.
Published: (2025)
TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
by: Malmsten, Emil, et al.
Published: (2025)
by: Malmsten, Emil, et al.
Published: (2025)
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
by: Oren, Yaniv, et al.
Published: (2026)
by: Oren, Yaniv, et al.
Published: (2026)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)
by: Murgoci, Vlad, et al.
Published: (2026)
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
by: de Vries, Joery A., et al.
Published: (2026)
by: de Vries, Joery A., et al.
Published: (2026)
Generalisation to unseen topologies: Towards control of biological neural network activity
by: Engwegen, Laurens, et al.
Published: (2024)
by: Engwegen, Laurens, et al.
Published: (2024)
Positive Experience Reflection for Agents in Interactive Text Environments
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
Trust-Region Twisted Policy Improvement
by: de Vries, Joery A., et al.
Published: (2025)
by: de Vries, Joery A., et al.
Published: (2025)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Modular Recurrence in Contextual MDPs for Universal Morphology Control
by: Engwegen, Laurens, et al.
Published: (2025)
by: Engwegen, Laurens, et al.
Published: (2025)
Uncertainty Quantification and Data Efficiency in AI: An Information-Theoretic Perspective
by: Simeone, Osvaldo, et al.
Published: (2025)
by: Simeone, Osvaldo, et al.
Published: (2025)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Quantifying Epistemic Uncertainty in Diffusion Models
by: Gupta, Aditi, et al.
Published: (2026)
by: Gupta, Aditi, et al.
Published: (2026)
Low-distortion and GPU-compatible Tree Embeddings in Hyperbolic Space
by: van Spengler, Max, et al.
Published: (2025)
by: van Spengler, Max, et al.
Published: (2025)
Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2025)
by: Vincent, Théo, et al.
Published: (2025)
Per-Domain Generalizing Policies: On Learning Efficient and Robust Q-Value Functions (Extended Version with Technical Appendix)
by: Müller, Nicola J., et al.
Published: (2026)
by: Müller, Nicola J., et al.
Published: (2026)
Hierarchical Universal Value Function Approximators
by: Arora, Rushiv
Published: (2024)
by: Arora, Rushiv
Published: (2024)
Improving Counterfactual Truthfulness for Molecular Property Prediction through Uncertainty Quantification
by: Teufel, Jonas, et al.
Published: (2025)
by: Teufel, Jonas, et al.
Published: (2025)
Online Learning under Haphazard Input Conditions: A Comprehensive Review and Analysis
by: Agarwal, Rohit, et al.
Published: (2024)
by: Agarwal, Rohit, et al.
Published: (2024)
packetLSTM: Dynamic LSTM Framework for Streaming Data with Varying Feature Space
by: Agarwal, Rohit, et al.
Published: (2024)
by: Agarwal, Rohit, et al.
Published: (2024)
Adversarial Attacks on Hyperbolic Networks
by: van Spengler, Max, et al.
Published: (2024)
by: van Spengler, Max, et al.
Published: (2024)
Active Timepoint Selection for Learning Measure-Valued Trajectories
by: Huynh, Nicolas, et al.
Published: (2026)
by: Huynh, Nicolas, et al.
Published: (2026)
Quasimetric Value Functions with Dense Rewards
by: Valieva, Khadichabonu, et al.
Published: (2024)
by: Valieva, Khadichabonu, et al.
Published: (2024)
Diverse Transformer Decoding for Offline Reinforcement Learning Using Financial Algorithmic Approaches
by: Elbaz, Dan, et al.
Published: (2025)
by: Elbaz, Dan, et al.
Published: (2025)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025)
by: Wagner, Moritz, et al.
Published: (2025)
Inferring Transition Dynamics from Value Functions
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
Graph Neural Networks for Transmission Grid Topology Control: Busbar Information Asymmetry and Heterogeneous Representations
by: de Jong, Matthijs, et al.
Published: (2025)
by: de Jong, Matthijs, et al.
Published: (2025)
Similar Items
-
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025) -
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024) -
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024) -
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025) -
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)