Regret Minimization via Saddle Point Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Kirschner, Johannes, Bakhtiari, Seyed Alireza, Chandak, Kushagra, Tkachuk, Volodymyr, Szepesvári, Csaba |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024)
by: Tkachuk, Volodymyr, et al.
Published: (2024)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Rectifying Regression in Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026)
by: Bakhtiari, Alireza, et al.
Published: (2026)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
Online Min-Max Optimization: From Individual Regrets to Cumulative Saddle Points
by: Vyas, Abhijeet, et al.
Published: (2026)
by: Vyas, Abhijeet, et al.
Published: (2026)
Sharp analysis of linear ensemble sampling
by: Akhavan, Arya, et al.
Published: (2026)
by: Akhavan, Arya, et al.
Published: (2026)
Exploration via linearly perturbed loss minimisation
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Balancing optimism and pessimism in offline-to-online learning
by: Sentenac, Flore, et al.
Published: (2025)
by: Sentenac, Flore, et al.
Published: (2025)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning
by: Schlegel, Matthew, et al.
Published: (2026)
by: Schlegel, Matthew, et al.
Published: (2026)
Federated Composite Saddle Point Optimization
by: Bai, Site, et al.
Published: (2023)
by: Bai, Site, et al.
Published: (2023)
Simultaneous Learning and Optimization via Misspecified Saddle Point Problems
by: Ahmadi, Mohammad Mahdi, et al.
Published: (2025)
by: Ahmadi, Mohammad Mahdi, et al.
Published: (2025)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
Quantization Avoids Saddle Points in Distributed Optimization
by: Bo, Yanan, et al.
Published: (2024)
by: Bo, Yanan, et al.
Published: (2024)
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
by: Singh, Rahul, et al.
Published: (2026)
by: Singh, Rahul, et al.
Published: (2026)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Principled Confidence Estimation for Deep Computed Tomography
by: Gätzner, Matteo, et al.
Published: (2026)
by: Gätzner, Matteo, et al.
Published: (2026)
Efficiently Escaping Saddle Points for Policy Optimization
by: Khorasani, Sadegh, et al.
Published: (2023)
by: Khorasani, Sadegh, et al.
Published: (2023)
Calibrated Regression Against An Adversary Without Regret
by: Deshpande, Shachi, et al.
Published: (2023)
by: Deshpande, Shachi, et al.
Published: (2023)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
A Saddle Point Remedy: Power of Variable Elimination in Non-convex Optimization
by: Gan, Min, et al.
Published: (2025)
by: Gan, Min, et al.
Published: (2025)
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Dimension-Free Saddle-Point Escape in Muon
by: Long, Yanlin, et al.
Published: (2026)
by: Long, Yanlin, et al.
Published: (2026)
Neural Network-based High-index Saddle Dynamics Method for Searching Saddle Points and Solution Landscape
by: Liu, Yuankai, et al.
Published: (2024)
by: Liu, Yuankai, et al.
Published: (2024)
Confidence Estimation via Sequential Likelihood Mixing
by: Kirschner, Johannes, et al.
Published: (2025)
by: Kirschner, Johannes, et al.
Published: (2025)
Proximal Point Method for Online Saddle Point Problem
by: Meng, Qing-xin, et al.
Published: (2024)
by: Meng, Qing-xin, et al.
Published: (2024)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
by: Bernasconi, Martino, et al.
Published: (2024)
by: Bernasconi, Martino, et al.
Published: (2024)
Optimal-Point Variance Reduction For Bayesian Optimization With Regret Guarantee
by: Takeno, Shion
Published: (2026)
by: Takeno, Shion
Published: (2026)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Hierarchical Deep Counterfactual Regret Minimization
by: Chen, Jiayu, et al.
Published: (2023)
by: Chen, Jiayu, et al.
Published: (2023)
Invariance-Based Dynamic Regret Minimization
by: Lazzaretto, Margherita, et al.
Published: (2026)
by: Lazzaretto, Margherita, et al.
Published: (2026)
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2025)
by: Genalti, Gianmarco, et al.
Published: (2025)
Similar Items
-
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024) -
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026) -
Rectifying Regression in Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025) -
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026) -
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)