Clustered KL-barycenter design for policy evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Weissmann, Simon, Freihaut, Till, Vernade, Claire, Ramponi, Giorgia, Döring, Leif |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
por: Freihaut, Till, et al.
Publicado: (2024)
por: Freihaut, Till, et al.
Publicado: (2024)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
por: Klein, Sara, et al.
Publicado: (2023)
por: Klein, Sara, et al.
Publicado: (2023)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
por: Kassing, Sebastian, et al.
Publicado: (2025)
por: Kassing, Sebastian, et al.
Publicado: (2025)
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Almost sure convergence rates of stochastic gradient methods under gradient domination
por: Weissmann, Simon, et al.
Publicado: (2024)
por: Weissmann, Simon, et al.
Publicado: (2024)
The Role of Target Update Frequencies in Q-Learning
por: Weissmann, Simon, et al.
Publicado: (2026)
por: Weissmann, Simon, et al.
Publicado: (2026)
Structure Matters: Dynamic Policy Gradient
por: Klein, Sara, et al.
Publicado: (2024)
por: Klein, Sara, et al.
Publicado: (2024)
Tight Sample Complexity Bounds for Entropic Best Policy Identification
por: Essakine, Amer, et al.
Publicado: (2026)
por: Essakine, Amer, et al.
Publicado: (2026)
Partially Observable Reinforcement Learning with Memory Traces
por: Eberhard, Onno, et al.
Publicado: (2025)
por: Eberhard, Onno, et al.
Publicado: (2025)
Non-Stationary Lipschitz Bandits
por: Nguyen, Nicolas, et al.
Publicado: (2025)
por: Nguyen, Nicolas, et al.
Publicado: (2025)
Commit to the Bit: Reactive Reinforcement Learning Done Right
por: Eberhard, Onno, et al.
Publicado: (2026)
por: Eberhard, Onno, et al.
Publicado: (2026)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
por: Macuglia, Maël, et al.
Publicado: (2025)
por: Macuglia, Maël, et al.
Publicado: (2025)
Random Function Descent
por: Benning, Felix, et al.
Publicado: (2023)
por: Benning, Felix, et al.
Publicado: (2023)
Variational Bayes Portfolio Construction
por: Nguyen, Nicolas, et al.
Publicado: (2024)
por: Nguyen, Nicolas, et al.
Publicado: (2024)
A Pontryagin Perspective on Reinforcement Learning
por: Eberhard, Onno, et al.
Publicado: (2024)
por: Eberhard, Onno, et al.
Publicado: (2024)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
por: Nguyen, Nicolas, et al.
Publicado: (2024)
por: Nguyen, Nicolas, et al.
Publicado: (2024)
Online Decision Deferral under Budget Constraints
por: Reid, Mirabel, et al.
Publicado: (2024)
por: Reid, Mirabel, et al.
Publicado: (2024)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
por: Balef, Amir Rezaei, et al.
Publicado: (2025)
por: Balef, Amir Rezaei, et al.
Publicado: (2025)
Preference Elicitation for Offline Reinforcement Learning
por: Pace, Alizée, et al.
Publicado: (2024)
por: Pace, Alizée, et al.
Publicado: (2024)
Quantization-Free Autoregressive Action Transformer
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2025)
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2025)
Truly No-Regret Learning in Constrained MDPs
por: Müller, Adrian, et al.
Publicado: (2024)
por: Müller, Adrian, et al.
Publicado: (2024)
Gradient Span Algorithms Make Predictable Progress in High Dimension
por: Benning, Felix, et al.
Publicado: (2024)
por: Benning, Felix, et al.
Publicado: (2024)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
por: Mutti, Mirco, et al.
Publicado: (2023)
por: Mutti, Mirco, et al.
Publicado: (2023)
On the Convergence of Single-Timescale Actor-Critic
por: Kumar, Navdeep, et al.
Publicado: (2024)
por: Kumar, Navdeep, et al.
Publicado: (2024)
Efficient Risk-sensitive Planning via Entropic Risk Measures
por: Marthe, Alexandre, et al.
Publicado: (2025)
por: Marthe, Alexandre, et al.
Publicado: (2025)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
por: Thoma, Vinzenz, et al.
Publicado: (2024)
por: Thoma, Vinzenz, et al.
Publicado: (2024)
Learning Acrobatic Flight from Preferences
por: Merk, Colin, et al.
Publicado: (2025)
por: Merk, Colin, et al.
Publicado: (2025)
An in depth look at the Procrustes-Wasserstein distance: properties and barycenters
por: Adamo, Davide, et al.
Publicado: (2025)
por: Adamo, Davide, et al.
Publicado: (2025)
An Approximate Ascent Approach To Prove Convergence of PPO
por: Doering, Leif, et al.
Publicado: (2026)
por: Doering, Leif, et al.
Publicado: (2026)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
por: Koren, Uri, et al.
Publicado: (2025)
por: Koren, Uri, et al.
Publicado: (2025)
ADDQ: Adaptive Distributional Double Q-Learning
por: Döring, Leif, et al.
Publicado: (2025)
por: Döring, Leif, et al.
Publicado: (2025)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
por: Kumar, Navdeep, et al.
Publicado: (2026)
por: Kumar, Navdeep, et al.
Publicado: (2026)
Beyond R-barycenters: an effective averaging method on Stiefel and Grassmann manifolds
por: Bouchard, Florent, et al.
Publicado: (2025)
por: Bouchard, Florent, et al.
Publicado: (2025)
Finite sample bounds for barycenter estimation in geodesic spaces
por: Brunel, Victor-Emmanuel, et al.
Publicado: (2025)
por: Brunel, Victor-Emmanuel, et al.
Publicado: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
por: Kumar, Navdeep, et al.
Publicado: (2025)
por: Kumar, Navdeep, et al.
Publicado: (2025)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
por: Baur, Raphaël, et al.
Publicado: (2026)
por: Baur, Raphaël, et al.
Publicado: (2026)
Adaptive Kernel Selection for Stein Variational Gradient Descent
por: Melcher, Moritz, et al.
Publicado: (2025)
por: Melcher, Moritz, et al.
Publicado: (2025)
Ejemplares similares
-
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
por: Freihaut, Till, et al.
Publicado: (2024) -
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025) -
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
por: Viano, Luca, et al.
Publicado: (2026) -
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
por: Klein, Sara, et al.
Publicado: (2023) -
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
por: Kassing, Sebastian, et al.
Publicado: (2025)