Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
Fuente:
arXiv
Guardado en:
| Autores principales: | Di, Qiwei, He, Jiafan, Zhou, Dongruo, Gu, Quanquan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
por: Di, Qiwei, et al.
Publicado: (2024)
por: Di, Qiwei, et al.
Publicado: (2024)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
por: He, Jiafan, et al.
Publicado: (2025)
por: He, Jiafan, et al.
Publicado: (2025)
Achieving Constant Regret in Linear Markov Decision Processes
por: Zhang, Weitong, et al.
Publicado: (2024)
por: Zhang, Weitong, et al.
Publicado: (2024)
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
por: Yu, Yue, et al.
Publicado: (2025)
por: Yu, Yue, et al.
Publicado: (2025)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
por: Zhao, Heyang, et al.
Publicado: (2023)
por: Zhao, Heyang, et al.
Publicado: (2023)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
por: Di, Qiwei, et al.
Publicado: (2023)
por: Di, Qiwei, et al.
Publicado: (2023)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
por: Zhang, Weitong, et al.
Publicado: (2021)
por: Zhang, Weitong, et al.
Publicado: (2021)
Best-of-Majority: Minimax-Optimal Strategy for Pass@$k$ Inference Scaling
por: Di, Qiwei, et al.
Publicado: (2025)
por: Di, Qiwei, et al.
Publicado: (2025)
Regret Guarantees for Linear Contextual Stochastic Shortest Path
por: Polikar, Dor, et al.
Publicado: (2025)
por: Polikar, Dor, et al.
Publicado: (2025)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
por: Di, Qiwei, et al.
Publicado: (2023)
por: Di, Qiwei, et al.
Publicado: (2023)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
por: Li, Long-Fei, et al.
Publicado: (2024)
por: Li, Long-Fei, et al.
Publicado: (2024)
Reinforcement Learning from Human Feedback with Active Queries
por: Ji, Kaixuan, et al.
Publicado: (2024)
por: Ji, Kaixuan, et al.
Publicado: (2024)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
por: Lee, Joongkyu, et al.
Publicado: (2024)
por: Lee, Joongkyu, et al.
Publicado: (2024)
Towards Robust Model-Based Reinforcement Learning Against Adversarial Corruption
por: Ye, Chenlu, et al.
Publicado: (2024)
por: Ye, Chenlu, et al.
Publicado: (2024)
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
por: Li, Shangzhe, et al.
Publicado: (2025)
por: Li, Shangzhe, et al.
Publicado: (2025)
Accelerated Preference Optimization for Large Language Model Alignment
por: He, Jiafan, et al.
Publicado: (2024)
por: He, Jiafan, et al.
Publicado: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
por: Zhang, Junkai, et al.
Publicado: (2023)
por: Zhang, Junkai, et al.
Publicado: (2023)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
por: Yang, Chen, et al.
Publicado: (2024)
por: Yang, Chen, et al.
Publicado: (2024)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
por: Zhang, Junkai, et al.
Publicado: (2024)
por: Zhang, Junkai, et al.
Publicado: (2024)
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
por: Li, Runjia, et al.
Publicado: (2024)
por: Li, Runjia, et al.
Publicado: (2024)
Regret Lower Bounds for Decentralized Multi-Agent Stochastic Shortest Path Problems
por: Chavan, Utkarsh U., et al.
Publicado: (2025)
por: Chavan, Utkarsh U., et al.
Publicado: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
por: Wang, Zhiyong, et al.
Publicado: (2024)
por: Wang, Zhiyong, et al.
Publicado: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
por: Zhou, Dongruo, et al.
Publicado: (2018)
por: Zhou, Dongruo, et al.
Publicado: (2018)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
por: Liu, Shuai, et al.
Publicado: (2026)
por: Liu, Shuai, et al.
Publicado: (2026)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
por: Zhang, Shiyuan, et al.
Publicado: (2026)
por: Zhang, Shiyuan, et al.
Publicado: (2026)
Near-Optimal Algorithms for Making the Gradient Small in Stochastic Minimax Optimization
por: Chen, Lesi, et al.
Publicado: (2022)
por: Chen, Lesi, et al.
Publicado: (2022)
Bridging Local and Global Knowledge: Cascaded Mixture-of-Experts Learning for Near-Shortest Path Routing
por: Chen, Yung-Fu, et al.
Publicado: (2026)
por: Chen, Yung-Fu, et al.
Publicado: (2026)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
por: Zhao, Heyang, et al.
Publicado: (2025)
por: Zhao, Heyang, et al.
Publicado: (2025)
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
por: Zhang, Haochen, et al.
Publicado: (2026)
por: Zhang, Haochen, et al.
Publicado: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
por: Cassel, Asaf, et al.
Publicado: (2024)
por: Cassel, Asaf, et al.
Publicado: (2024)
Convergent Reinforcement Learning Algorithms for Stochastic Shortest Path Problem
por: Guin, Soumyajit, et al.
Publicado: (2025)
por: Guin, Soumyajit, et al.
Publicado: (2025)
Minimax-Optimal Policy Regret in Partially Observable Markov Games
por: Arora, Raman
Publicado: (2026)
por: Arora, Raman
Publicado: (2026)
Stochastic Shortest Path with Sparse Adversarial Costs
por: Johnson, Emmeran, et al.
Publicado: (2025)
por: Johnson, Emmeran, et al.
Publicado: (2025)
Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits
por: Ye, Zichun, et al.
Publicado: (2025)
por: Ye, Zichun, et al.
Publicado: (2025)
Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits
por: Liu, Jingyu, et al.
Publicado: (2025)
por: Liu, Jingyu, et al.
Publicado: (2025)
Minimax-Optimal Spectral Clustering with Covariance Projection for High-Dimensional Anisotropic Mixtures
por: Huang, Chengzhu, et al.
Publicado: (2025)
por: Huang, Chengzhu, et al.
Publicado: (2025)
Knowledge-Guided Machine Learning for Stabilizing Near-Shortest Path Routing
por: Chen, Yung-Fu, et al.
Publicado: (2025)
por: Chen, Yung-Fu, et al.
Publicado: (2025)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
por: Rowland, Mark, et al.
Publicado: (2024)
por: Rowland, Mark, et al.
Publicado: (2024)
Ejemplares similares
-
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
por: Di, Qiwei, et al.
Publicado: (2024) -
Variance-Dependent Regret Lower Bounds for Contextual Bandits
por: He, Jiafan, et al.
Publicado: (2025) -
Achieving Constant Regret in Linear Markov Decision Processes
por: Zhang, Weitong, et al.
Publicado: (2024) -
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
por: Yu, Yue, et al.
Publicado: (2025) -
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
por: Zhao, Heyang, et al.
Publicado: (2023)