Augmented Bayesian Policy Search
Fuente:
arXiv
Saved in:
| Main Authors: | Kallel, Mahdi, Basu, Debabrota, Akrour, Riad, D'Eramo, Carlo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement
by: Kallel, Mahdi, et al.
Published: (2026)
by: Kallel, Mahdi, et al.
Published: (2026)
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
by: Hendawy, Ahmed, et al.
Published: (2025)
by: Hendawy, Ahmed, et al.
Published: (2025)
Deterministic Exploration via Stationary Bellman Error Maximization
by: Griesbach, Sebastian, et al.
Published: (2024)
by: Griesbach, Sebastian, et al.
Published: (2024)
Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization
by: Griesbach, Sebastian, et al.
Published: (2025)
by: Griesbach, Sebastian, et al.
Published: (2025)
Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts
by: Hendawy, Ahmed, et al.
Published: (2023)
by: Hendawy, Ahmed, et al.
Published: (2023)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
StaQ it! Growing neural networks for Policy Mirror Descent
by: Shilova, Alena, et al.
Published: (2025)
by: Shilova, Alena, et al.
Published: (2025)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
by: Kohler, Hector, et al.
Published: (2023)
by: Kohler, Hector, et al.
Published: (2023)
On the Benefit of Optimal Transport for Curriculum Reinforcement Learning
by: Klink, Pascal, et al.
Published: (2023)
by: Klink, Pascal, et al.
Published: (2023)
Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning
by: Farr, Noah, et al.
Published: (2026)
by: Farr, Noah, et al.
Published: (2026)
Dynamic Obstacle Avoidance with Bounded Rationality Adversarial Reinforcement Learning
by: Holgado-Alvarez, Jose-Luis, et al.
Published: (2025)
by: Holgado-Alvarez, Jose-Luis, et al.
Published: (2025)
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
by: Driss, Brahim, et al.
Published: (2025)
by: Driss, Brahim, et al.
Published: (2025)
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs
by: Kohler, Hector, et al.
Published: (2025)
by: Kohler, Hector, et al.
Published: (2025)
Sharing Knowledge in Multi-Task Deep Reinforcement Learning
by: D'Eramo, Carlo, et al.
Published: (2024)
by: D'Eramo, Carlo, et al.
Published: (2024)
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2025)
by: Vincent, Théo, et al.
Published: (2025)
Preference-based Pure Exploration
by: Shukla, Apurv, et al.
Published: (2024)
by: Shukla, Apurv, et al.
Published: (2024)
Domain Randomization via Entropy Maximization
by: Tiboni, Gabriele, et al.
Published: (2023)
by: Tiboni, Gabriele, et al.
Published: (2023)
Stochastic Online Instrumental Variable Regression: Regrets for Endogeneity and Bandit Feedback
by: Della Vecchia, Riccardo, et al.
Published: (2023)
by: Della Vecchia, Riccardo, et al.
Published: (2023)
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints
by: Das, Udvas, et al.
Published: (2024)
by: Das, Udvas, et al.
Published: (2024)
Some Targets Are Harder to Identify than Others: Quantifying the Target-dependent Membership Leakage
by: Azize, Achraf, et al.
Published: (2024)
by: Azize, Achraf, et al.
Published: (2024)
Auditing Fairness under Model Updates: Fundamental Complexity and Property-Preserving Updates
by: Ajarra, Ayoub, et al.
Published: (2026)
by: Ajarra, Ayoub, et al.
Published: (2026)
Parameterized Projected Bellman Operator
by: Vincent, Théo, et al.
Published: (2023)
by: Vincent, Théo, et al.
Published: (2023)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Isoperimetry is All We Need: Langevin Posterior Sampling for RL with Sublinear Regret
by: Jorge, Emilio, et al.
Published: (2024)
by: Jorge, Emilio, et al.
Published: (2024)
Sublinear Algorithms for Wasserstein and Total Variation Distances: Applications to Fairness and Privacy Auditing
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Concentrated Differential Privacy for Bandits
by: Azize, Achraf, et al.
Published: (2023)
by: Azize, Achraf, et al.
Published: (2023)
FLIPHAT: Joint Differential Privacy for High Dimensional Sparse Linear Bandits
by: Chakraborty, Sunrit, et al.
Published: (2024)
by: Chakraborty, Sunrit, et al.
Published: (2024)
Pure Exploration in Bandits with Linear Constraints
by: Carlsson, Emil, et al.
Published: (2023)
by: Carlsson, Emil, et al.
Published: (2023)
Active Fourier Auditor for Estimating Distributional Properties of ML Models
by: Ajarra, Ayoub, et al.
Published: (2024)
by: Ajarra, Ayoub, et al.
Published: (2024)
Sequential Membership Inference Attacks
by: Michel, Thomas, et al.
Published: (2026)
by: Michel, Thomas, et al.
Published: (2026)
DP-SPRT: Differentially Private Sequential Probability Ratio Tests
by: Michel, Thomas, et al.
Published: (2025)
by: Michel, Thomas, et al.
Published: (2025)
Gradient Iterated Temporal-Difference Learning
by: Vincent, Théo, et al.
Published: (2026)
by: Vincent, Théo, et al.
Published: (2026)
Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2025)
by: Vincent, Théo, et al.
Published: (2025)
Lagrangian-based Equilibrium Propagation: generalisation to arbitrary boundary conditions & equivalence with Hamiltonian Echo Learning
by: Pourcel, Guillaume, et al.
Published: (2025)
by: Pourcel, Guillaume, et al.
Published: (2025)
Light to Heavy, Brief to Eternal: An Axion for Every Occasion (in the Early Universe)
by: D'Eramo, Francesco
Published: (2026)
by: D'Eramo, Francesco
Published: (2026)
FraPPE: Fast and Efficient Preference-based Pure Exploration
by: Das, Udvas, et al.
Published: (2025)
by: Das, Udvas, et al.
Published: (2025)
Similar Items
-
Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement
by: Kallel, Mahdi, et al.
Published: (2026) -
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
by: Hendawy, Ahmed, et al.
Published: (2025) -
Deterministic Exploration via Stationary Bellman Error Maximization
by: Griesbach, Sebastian, et al.
Published: (2024) -
Learning to Explore in Diverse Reward Settings via Temporal-Difference-Error Maximization
by: Griesbach, Sebastian, et al.
Published: (2025) -
Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts
by: Hendawy, Ahmed, et al.
Published: (2023)