A Diffusion Analysis of Policy Gradient for Stochastic Bandits
Fuente:
arXiv
Saved in:
| Main Author: | Lattimore, Tor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Stein-Rule Shrinkage for Stochastic Gradient Estimation in High Dimensions
by: Arashi, M., et al.
Published: (2026)
by: Arashi, M., et al.
Published: (2026)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
by: Roy, Saptarshi, et al.
Published: (2026)
by: Roy, Saptarshi, et al.
Published: (2026)
Bandit Convex Optimisation
by: Lattimore, Tor
Published: (2024)
by: Lattimore, Tor
Published: (2024)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning
by: Cao, Junyu, et al.
Published: (2026)
by: Cao, Junyu, et al.
Published: (2026)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Learning with Differentially Private (Sliced) Wasserstein Gradients
by: Rodríguez-Vítores, David, et al.
Published: (2025)
by: Rodríguez-Vítores, David, et al.
Published: (2025)
Conformal Policy Control
by: Prinster, Drew, et al.
Published: (2026)
by: Prinster, Drew, et al.
Published: (2026)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Diffusion Posterior Sampling is Computationally Intractable
by: Gupta, Shivam, et al.
Published: (2024)
by: Gupta, Shivam, et al.
Published: (2024)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
by: Farghly, Tyler, et al.
Published: (2025)
by: Farghly, Tyler, et al.
Published: (2025)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
by: Tohme, Tony, et al.
Published: (2023)
by: Tohme, Tony, et al.
Published: (2023)
Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2026)
by: Chakraborty, Saptarshi, et al.
Published: (2026)
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
by: Mei, Song
Published: (2024)
by: Mei, Song
Published: (2024)
Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression
by: Liang, Haodong, et al.
Published: (2025)
by: Liang, Haodong, et al.
Published: (2025)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
A Simple and Optimal Policy Design with Safety against Heavy-Tailed Risk for Stochastic Bandits
by: Simchi-Levi, David, et al.
Published: (2022)
by: Simchi-Levi, David, et al.
Published: (2022)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
A Score-Based Density Formula, with Applications in Diffusion Generative Models
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Identifying All ε-Best Arms in (Misspecified) Linear Bandits
by: Li, Zhekai, et al.
Published: (2025)
by: Li, Zhekai, et al.
Published: (2025)
Interaction Testing in Variation Analysis
by: Plecko, Drago
Published: (2024)
by: Plecko, Drago
Published: (2024)
Truncated LinUCB for Stochastic Linear Bandits
by: Song, Yanglei, et al.
Published: (2022)
by: Song, Yanglei, et al.
Published: (2022)
On the Separability of Information in Diffusion Models
by: Premkumar, Akhil
Published: (2025)
by: Premkumar, Akhil
Published: (2025)
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
by: Bae, Youngkyoung, et al.
Published: (2024)
by: Bae, Youngkyoung, et al.
Published: (2024)
Statistical inference with belief functions: A survey
by: Cuzzolin, Fabio
Published: (2026)
by: Cuzzolin, Fabio
Published: (2026)
A Quantitative Characterization of Forgetting in Post-Training
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
by: Kontorovich, Aryeh, et al.
Published: (2026)
by: Kontorovich, Aryeh, et al.
Published: (2026)
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Deep Ensembles for Epistemic Uncertainty: A Frequentist Perspective
by: Jain, Anchit, et al.
Published: (2025)
by: Jain, Anchit, et al.
Published: (2025)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
by: Yan, Hedong
Published: (2025)
by: Yan, Hedong
Published: (2025)
Similar Items
-
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026) -
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025) -
Stein-Rule Shrinkage for Stochastic Gradient Estimation in High Dimensions
by: Arashi, M., et al.
Published: (2026) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026) -
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
by: Roy, Saptarshi, et al.
Published: (2026)