Convergence of a L2 regularized Policy Gradient Algorithm for the Multi Armed Bandit
Fuente:
arXiv
Guardado en:
| Autores principales: | Anita, Stefana, Turinici, Gabriel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Introduction to Multi-Armed Bandits
por: Slivkins, Aleksandrs
Publicado: (2019)
por: Slivkins, Aleksandrs
Publicado: (2019)
Algorithms and data structures for automatic precision estimation of neural networks
por: Netay, Igor V.
Publicado: (2025)
por: Netay, Igor V.
Publicado: (2025)
On the Robustness of the Successive Projection Algorithm
por: Barbarino, Giovanni, et al.
Publicado: (2024)
por: Barbarino, Giovanni, et al.
Publicado: (2024)
Sinkhorn Algorithm for Sequentially Composed Optimal Transports
por: Watanabe, Kazuki, et al.
Publicado: (2024)
por: Watanabe, Kazuki, et al.
Publicado: (2024)
Stochastic Multi-Objective Multi-Armed Bandits: Regret Definition and Algorithm
por: Davoodi, Mansoor, et al.
Publicado: (2025)
por: Davoodi, Mansoor, et al.
Publicado: (2025)
Randomized Kaczmarz Methods with Beyond-Krylov Convergence
por: Dereziński, Michał, et al.
Publicado: (2025)
por: Dereziński, Michał, et al.
Publicado: (2025)
Towards Universal Convergence of Backward Error in Linear System Solvers
por: Dereziński, Michał, et al.
Publicado: (2026)
por: Dereziński, Michał, et al.
Publicado: (2026)
Adversarial Attacks on Combinatorial Multi-Armed Bandits
por: Balasubramanian, Rishab, et al.
Publicado: (2023)
por: Balasubramanian, Rishab, et al.
Publicado: (2023)
Unlearning Offline Stochastic Multi-Armed Bandits
por: Ye, Zichun, et al.
Publicado: (2026)
por: Ye, Zichun, et al.
Publicado: (2026)
Distributed Least Squares in Small Space via Sketching and Bias Reduction
por: Garg, Sachin, et al.
Publicado: (2024)
por: Garg, Sachin, et al.
Publicado: (2024)
Black-Box $k$-to-$1$-PCA Reductions: Theory and Applications
por: Jambulapati, Arun, et al.
Publicado: (2024)
por: Jambulapati, Arun, et al.
Publicado: (2024)
Optimal Embedding Dimension for Sparse Subspace Embeddings
por: Chenakkod, Shabarish, et al.
Publicado: (2023)
por: Chenakkod, Shabarish, et al.
Publicado: (2023)
Accelerating Power Method with Fast Sketching for Stronger Low-Rank Approximation
por: Chenakkod, Shabarish, et al.
Publicado: (2026)
por: Chenakkod, Shabarish, et al.
Publicado: (2026)
Query Efficient Structured Matrix Learning
por: Amsel, Noah, et al.
Publicado: (2025)
por: Amsel, Noah, et al.
Publicado: (2025)
Arithmetical Binary Decision Tree Traversals
por: Zhang, Jinxiong
Publicado: (2022)
por: Zhang, Jinxiong
Publicado: (2022)
Algorithmic warm starts for Hamiltonian Monte Carlo
por: Zhang, Matthew S., et al.
Publicado: (2026)
por: Zhang, Matthew S., et al.
Publicado: (2026)
Fine-grained Analysis and Faster Algorithms for Iteratively Solving Linear Systems
por: Dereziński, Michał, et al.
Publicado: (2024)
por: Dereziński, Michał, et al.
Publicado: (2024)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
por: Pedramfar, Mohammad, et al.
Publicado: (2023)
por: Pedramfar, Mohammad, et al.
Publicado: (2023)
Nearly-tight Approximation Guarantees for the Improving Multi-Armed Bandits Problem
por: Blum, Avrim, et al.
Publicado: (2024)
por: Blum, Avrim, et al.
Publicado: (2024)
Vanishing L2 regularization for the softmax Multi Armed Bandit
por: Anita, Stefana-Lucia, et al.
Publicado: (2026)
por: Anita, Stefana-Lucia, et al.
Publicado: (2026)
Optimal Oblivious Subspace Embeddings with Near-optimal Sparsity
por: Chenakkod, Shabarish, et al.
Publicado: (2024)
por: Chenakkod, Shabarish, et al.
Publicado: (2024)
Well-Conditioned Oblivious Perturbations in Linear Space
por: Chenakkod, Shabarish, et al.
Publicado: (2026)
por: Chenakkod, Shabarish, et al.
Publicado: (2026)
Optimal Subspace Embeddings: Resolving Nelson-Nguyen Conjecture Up to Sub-Polylogarithmic Factors
por: Chenakkod, Shabarish, et al.
Publicado: (2025)
por: Chenakkod, Shabarish, et al.
Publicado: (2025)
Faster Linear Systems and Matrix Norm Approximation via Multi-level Sketched Preconditioning
por: Dereziński, Michał, et al.
Publicado: (2024)
por: Dereziński, Michał, et al.
Publicado: (2024)
Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits
por: He, Yuchen, et al.
Publicado: (2024)
por: He, Yuchen, et al.
Publicado: (2024)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
por: Turinici, Gabriel
Publicado: (2026)
por: Turinici, Gabriel
Publicado: (2026)
Analysis of Different Algorithmic Design Techniques for Seam Carving
por: Aijaz, Owais, et al.
Publicado: (2024)
por: Aijaz, Owais, et al.
Publicado: (2024)
Pareto Optimal Algorithmic Recourse in Multi-cost Function
por: Chen, Wen-Ling, et al.
Publicado: (2025)
por: Chen, Wen-Ling, et al.
Publicado: (2025)
Algorithms and data structures for numerical computations with automatic precision estimation
por: Netay, Igor V.
Publicado: (2024)
por: Netay, Igor V.
Publicado: (2024)
Shifted Composition III: Local Error Framework for KL Divergence
por: Altschuler, Jason M., et al.
Publicado: (2024)
por: Altschuler, Jason M., et al.
Publicado: (2024)
Approaching Optimality for Solving Dense Linear Systems with Low-Rank Structure
por: Dereziński, Michał, et al.
Publicado: (2025)
por: Dereziński, Michał, et al.
Publicado: (2025)
Solving Dense Linear Systems Faster Than via Preconditioning
por: Dereziński, Michał, et al.
Publicado: (2023)
por: Dereziński, Michał, et al.
Publicado: (2023)
Iterative Refinement for $\ell_p$-norm Regression
por: Adil, Deeksha, et al.
Publicado: (2019)
por: Adil, Deeksha, et al.
Publicado: (2019)
On computing and the complexity of computing higher-order $U$-statistics, exactly
por: Chen, Xingyu, et al.
Publicado: (2025)
por: Chen, Xingyu, et al.
Publicado: (2025)
Lower bounds for trace estimation via Block Krylov and other methods
por: Yu, Shi Jie
Publicado: (2025)
por: Yu, Shi Jie
Publicado: (2025)
Tight Gap-Dependent Memory-Regret Trade-Off for Single-Pass Streaming Stochastic Multi-Armed Bandits
por: Ye, Zichun, et al.
Publicado: (2025)
por: Ye, Zichun, et al.
Publicado: (2025)
Stochastic Rounding 2.0, with a View towards Complexity Analysis
por: Drineas, Petros, et al.
Publicado: (2024)
por: Drineas, Petros, et al.
Publicado: (2024)
Fast EXP3 Algorithms
por: Sato, Ryoma, et al.
Publicado: (2025)
por: Sato, Ryoma, et al.
Publicado: (2025)
On Tradeoffs in Learning-Augmented Algorithms
por: Benomar, Ziyad, et al.
Publicado: (2025)
por: Benomar, Ziyad, et al.
Publicado: (2025)
Simulation of Graph Algorithms with Looped Transformers
por: de Luca, Artur Back, et al.
Publicado: (2024)
por: de Luca, Artur Back, et al.
Publicado: (2024)
Ejemplares similares
-
Introduction to Multi-Armed Bandits
por: Slivkins, Aleksandrs
Publicado: (2019) -
Algorithms and data structures for automatic precision estimation of neural networks
por: Netay, Igor V.
Publicado: (2025) -
On the Robustness of the Successive Projection Algorithm
por: Barbarino, Giovanni, et al.
Publicado: (2024) -
Sinkhorn Algorithm for Sequentially Composed Optimal Transports
por: Watanabe, Kazuki, et al.
Publicado: (2024) -
Stochastic Multi-Objective Multi-Armed Bandits: Regret Definition and Algorithm
por: Davoodi, Mansoor, et al.
Publicado: (2025)