Bayesian policy gradient and actor-critic algorithms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghavamzadeh, Mohammad, Engel, Yaakov, Valko, Michal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
von: Tang, Yunhao, et al.
Veröffentlicht: (2025)
Bandits on graphs and structures
von: Valko, Michal
Veröffentlicht: (2026)
von: Valko, Michal
Veröffentlicht: (2026)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
A single algorithm for both restless and rested rotting bandits
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
von: Seznec, Julien, et al.
Veröffentlicht: (2026)
Feature importance analysis for patient management decisions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
Distance metric learning for conditional anomaly detection
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Learning from a single labeled face and a stream of unlabeled data
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
Revealing graph bandits for maximizing local influence
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Extreme bandits
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
von: Carpentier, Alexandra, et al.
Veröffentlicht: (2026)
Contextual Bandits with Stage-wise Constraints
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
Conservative Contextual Bandits: Beyond Linear Representations
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
Active multiple matrix completion with adaptive confidence sets
von: Locatelli, Andrea, et al.
Veröffentlicht: (2026)
von: Locatelli, Andrea, et al.
Veröffentlicht: (2026)
Bandits attack function optimization
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
von: Preux, Philippe, et al.
Veröffentlicht: (2026)
Stochastic simultaneous optimistic optimization
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Online learning with Erdős-Rényi side-observation graphs
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Large-scale semi-supervised learning with online spectral graph sparsification
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Adaptive multi-fidelity optimization with fast learning rates
von: Fiegel, Come, et al.
Veröffentlicht: (2026)
von: Fiegel, Come, et al.
Veröffentlicht: (2026)
Analysis of Nystrom method with sequential ridge leverage scores
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Online learning with noisy side observations
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Pack only the essentials: Adaptive dictionary learning for kernel ridge regression
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
von: Calandriello, Daniele, et al.
Veröffentlicht: (2026)
Learning predictive models for combinations of heterogeneous proteomic data sources
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Language Generation with Replay: A Learning-Theoretic View of Model Collapse
von: Racca, Giorgio, et al.
Veröffentlicht: (2026)
von: Racca, Giorgio, et al.
Veröffentlicht: (2026)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
Generalizing soft actor-critic algorithms to discrete action spaces
von: Zhang, Le, et al.
Veröffentlicht: (2024)
von: Zhang, Le, et al.
Veröffentlicht: (2024)
Black-box optimization of noisy functions with unknown smoothness
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
von: Grill, Jean-Bastien, et al.
Veröffentlicht: (2026)
On two ways to use determinantal point processes for Monte Carlo integration
von: Gautier, Guillaume, et al.
Veröffentlicht: (2026)
von: Gautier, Guillaume, et al.
Veröffentlicht: (2026)
Adaptive Bayesian Optimization for Robust Identification of Stochastic Dynamical Systems
von: Xu, Jinwen, et al.
Veröffentlicht: (2025)
von: Xu, Jinwen, et al.
Veröffentlicht: (2025)
Diffusion Policy with Bayesian Expert Selection for Active Multi-Target Tracking
von: Xiang, Haotian, et al.
Veröffentlicht: (2026)
von: Xiang, Haotian, et al.
Veröffentlicht: (2026)
Spectral Thompson sampling
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
Semi-supervised learning with max-margin graph cuts
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
von: Tarbouriech, Jean, et al.
Veröffentlicht: (2026)
von: Tarbouriech, Jean, et al.
Veröffentlicht: (2026)
Spectral bandits for smooth graph functions
von: Valko, Michal, et al.
Veröffentlicht: (2026)
von: Valko, Michal, et al.
Veröffentlicht: (2026)
Efficient learning by implicit exploration in bandit problems with side observations
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
von: Kocak, Tomas, et al.
Veröffentlicht: (2026)
Budgeted Online Influence Maximization
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
von: Perrault, Pierre, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026) -
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
von: Valko, Michal
Veröffentlicht: (2026) -
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023) -
RL-finetuning LLMs from on- and off-policy data with a single algorithm
von: Tang, Yunhao, et al.
Veröffentlicht: (2025) -
Bandits on graphs and structures
von: Valko, Michal
Veröffentlicht: (2026)