Clustered KL-barycenter design for policy evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Weissmann, Simon, Freihaut, Till, Vernade, Claire, Ramponi, Giorgia, Döring, Leif |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
von: Klein, Sara, et al.
Veröffentlicht: (2023)
von: Klein, Sara, et al.
Veröffentlicht: (2023)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
Rate optimal learning of equilibria from data
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
Almost sure convergence rates of stochastic gradient methods under gradient domination
von: Weissmann, Simon, et al.
Veröffentlicht: (2024)
von: Weissmann, Simon, et al.
Veröffentlicht: (2024)
The Role of Target Update Frequencies in Q-Learning
von: Weissmann, Simon, et al.
Veröffentlicht: (2026)
von: Weissmann, Simon, et al.
Veröffentlicht: (2026)
Structure Matters: Dynamic Policy Gradient
von: Klein, Sara, et al.
Veröffentlicht: (2024)
von: Klein, Sara, et al.
Veröffentlicht: (2024)
Tight Sample Complexity Bounds for Entropic Best Policy Identification
von: Essakine, Amer, et al.
Veröffentlicht: (2026)
von: Essakine, Amer, et al.
Veröffentlicht: (2026)
Partially Observable Reinforcement Learning with Memory Traces
von: Eberhard, Onno, et al.
Veröffentlicht: (2025)
von: Eberhard, Onno, et al.
Veröffentlicht: (2025)
Non-Stationary Lipschitz Bandits
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2025)
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2025)
Commit to the Bit: Reactive Reinforcement Learning Done Right
von: Eberhard, Onno, et al.
Veröffentlicht: (2026)
von: Eberhard, Onno, et al.
Veröffentlicht: (2026)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
von: Macuglia, Maël, et al.
Veröffentlicht: (2025)
von: Macuglia, Maël, et al.
Veröffentlicht: (2025)
Random Function Descent
von: Benning, Felix, et al.
Veröffentlicht: (2023)
von: Benning, Felix, et al.
Veröffentlicht: (2023)
Variational Bayes Portfolio Construction
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2024)
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2024)
A Pontryagin Perspective on Reinforcement Learning
von: Eberhard, Onno, et al.
Veröffentlicht: (2024)
von: Eberhard, Onno, et al.
Veröffentlicht: (2024)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2024)
von: Nguyen, Nicolas, et al.
Veröffentlicht: (2024)
Online Decision Deferral under Budget Constraints
von: Reid, Mirabel, et al.
Veröffentlicht: (2024)
von: Reid, Mirabel, et al.
Veröffentlicht: (2024)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
Preference Elicitation for Offline Reinforcement Learning
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
Quantization-Free Autoregressive Action Transformer
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2025)
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2025)
Truly No-Regret Learning in Constrained MDPs
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
Gradient Span Algorithms Make Predictable Progress in High Dimension
von: Benning, Felix, et al.
Veröffentlicht: (2024)
von: Benning, Felix, et al.
Veröffentlicht: (2024)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Efficient Risk-sensitive Planning via Entropic Risk Measures
von: Marthe, Alexandre, et al.
Veröffentlicht: (2025)
von: Marthe, Alexandre, et al.
Veröffentlicht: (2025)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
von: Thoma, Vinzenz, et al.
Veröffentlicht: (2024)
von: Thoma, Vinzenz, et al.
Veröffentlicht: (2024)
Learning Acrobatic Flight from Preferences
von: Merk, Colin, et al.
Veröffentlicht: (2025)
von: Merk, Colin, et al.
Veröffentlicht: (2025)
An in depth look at the Procrustes-Wasserstein distance: properties and barycenters
von: Adamo, Davide, et al.
Veröffentlicht: (2025)
von: Adamo, Davide, et al.
Veröffentlicht: (2025)
An Approximate Ascent Approach To Prove Convergence of PPO
von: Doering, Leif, et al.
Veröffentlicht: (2026)
von: Doering, Leif, et al.
Veröffentlicht: (2026)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025)
von: Koren, Uri, et al.
Veröffentlicht: (2025)
ADDQ: Adaptive Distributional Double Q-Learning
von: Döring, Leif, et al.
Veröffentlicht: (2025)
von: Döring, Leif, et al.
Veröffentlicht: (2025)
Beyond R-barycenters: an effective averaging method on Stiefel and Grassmann manifolds
von: Bouchard, Florent, et al.
Veröffentlicht: (2025)
von: Bouchard, Florent, et al.
Veröffentlicht: (2025)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
Finite sample bounds for barycenter estimation in geodesic spaces
von: Brunel, Victor-Emmanuel, et al.
Veröffentlicht: (2025)
von: Brunel, Victor-Emmanuel, et al.
Veröffentlicht: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
Adaptive Kernel Selection for Stein Variational Gradient Descent
von: Melcher, Moritz, et al.
Veröffentlicht: (2025)
von: Melcher, Moritz, et al.
Veröffentlicht: (2025)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
von: Baur, Raphaël, et al.
Veröffentlicht: (2026)
von: Baur, Raphaël, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2024) -
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2025) -
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
von: Viano, Luca, et al.
Veröffentlicht: (2026) -
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
von: Klein, Sara, et al.
Veröffentlicht: (2023) -
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)