Best of Both Worlds: Regret Minimization versus Minimax Play
Fuente:
arXiv
Guardado en:
| Autores principales: | Müller, Adrian, Schneider, Jon, Skoulakis, Stratis, Viano, Luca, Cevher, Volkan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)
por: Viano, Luca, et al.
Publicado: (2024)
Polynomial Convergence of Bandit No-Regret Dynamics in Congestion Games
por: Dadi, Leello, et al.
Publicado: (2024)
por: Dadi, Leello, et al.
Publicado: (2024)
Efficient Continual Finite-Sum Minimization
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
Learning to Remove Cuts in Integer Linear Programming
por: Puigdemont, Pol, et al.
Publicado: (2024)
por: Puigdemont, Pol, et al.
Publicado: (2024)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
por: Viel, Stefano, et al.
Publicado: (2025)
por: Viel, Stefano, et al.
Publicado: (2025)
Continuous-Time Analysis of Heavy Ball Momentum in Min-Max Games
por: Feng, Yi, et al.
Publicado: (2025)
por: Feng, Yi, et al.
Publicado: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
por: Barla, Adam, et al.
Publicado: (2026)
por: Barla, Adam, et al.
Publicado: (2026)
Truly No-Regret Learning in Constrained MDPs
por: Müller, Adrian, et al.
Publicado: (2024)
por: Müller, Adrian, et al.
Publicado: (2024)
Optimism Without Regularization: Constant Regret in Zero-Sum Games
por: Lazarsfeld, John, et al.
Publicado: (2025)
por: Lazarsfeld, John, et al.
Publicado: (2025)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
por: Schlisselberg, Ofir, et al.
Publicado: (2025)
por: Schlisselberg, Ofir, et al.
Publicado: (2025)
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
SAMPa: Sharpness-aware Minimization Parallelized
por: Xie, Wanyun, et al.
Publicado: (2024)
por: Xie, Wanyun, et al.
Publicado: (2024)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
A Simple and Adaptive Learning Rate for FTRL in Online Learning with Minimax Regret of $Θ(T^{2/3})$ and its Application to Best-of-Both-Worlds
por: Tsuchiya, Taira, et al.
Publicado: (2024)
por: Tsuchiya, Taira, et al.
Publicado: (2024)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
por: Morales-Brotons, Daniel, et al.
Publicado: (2024)
por: Morales-Brotons, Daniel, et al.
Publicado: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
por: Wu, Yongtao, et al.
Publicado: (2025)
por: Wu, Yongtao, et al.
Publicado: (2025)
Near-optimal Swap Regret Minimization for Convex Losses
por: Hu, Lunjia, et al.
Publicado: (2026)
por: Hu, Lunjia, et al.
Publicado: (2026)
Best Arm Identification with Minimal Regret
por: Yang, Junwen, et al.
Publicado: (2024)
por: Yang, Junwen, et al.
Publicado: (2024)
Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits
por: Li, Mengmeng, et al.
Publicado: (2025)
por: Li, Mengmeng, et al.
Publicado: (2025)
Swap Regret Minimization Through Response-Based Approachability
por: Anagnostides, Ioannis, et al.
Publicado: (2026)
por: Anagnostides, Ioannis, et al.
Publicado: (2026)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
por: Haas, Moritz, et al.
Publicado: (2024)
por: Haas, Moritz, et al.
Publicado: (2024)
Minimax Optimal Simple Regret in Two-Armed Best-Arm Identification
por: Kato, Masahiro
Publicado: (2024)
por: Kato, Masahiro
Publicado: (2024)
Next-Token Prediction and Regret Minimization
por: Mohri, Mehryar, et al.
Publicado: (2026)
por: Mohri, Mehryar, et al.
Publicado: (2026)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
por: Liu, Fanghui, et al.
Publicado: (2024)
por: Liu, Fanghui, et al.
Publicado: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
por: Xie, Wanyun, et al.
Publicado: (2026)
por: Xie, Wanyun, et al.
Publicado: (2026)
Meta-Learning in Self-Play Regret Minimization
por: Sychrovský, David, et al.
Publicado: (2025)
por: Sychrovský, David, et al.
Publicado: (2025)
Best-of-Both-Worlds Algorithms for Linear Contextual Bandits
por: Kuroki, Yuko, et al.
Publicado: (2023)
por: Kuroki, Yuko, et al.
Publicado: (2023)
The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection
por: Zhang, Qingyang, et al.
Publicado: (2024)
por: Zhang, Qingyang, et al.
Publicado: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
por: Xie, Wanyun, et al.
Publicado: (2025)
por: Xie, Wanyun, et al.
Publicado: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
por: Erdogan, Mete, et al.
Publicado: (2025)
por: Erdogan, Mete, et al.
Publicado: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
por: Pethick, Thomas, et al.
Publicado: (2023)
por: Pethick, Thomas, et al.
Publicado: (2023)
Best-of-Both Worlds for linear contextual bandits with paid observations
por: Boyer, Nathan, et al.
Publicado: (2025)
por: Boyer, Nathan, et al.
Publicado: (2025)
Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes
por: Chen, Yu, et al.
Publicado: (2026)
por: Chen, Yu, et al.
Publicado: (2026)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
Mitigating Goal Misgeneralization via Minimax Regret
por: Sadek, Karim Abdel, et al.
Publicado: (2025)
por: Sadek, Karim Abdel, et al.
Publicado: (2025)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
por: Afzal, Arshia, et al.
Publicado: (2025)
por: Afzal, Arshia, et al.
Publicado: (2025)
Ejemplares similares
-
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024) -
Polynomial Convergence of Bandit No-Regret Dynamics in Congestion Games
por: Dadi, Leello, et al.
Publicado: (2024) -
Efficient Continual Finite-Sum Minimization
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024) -
Learning to Remove Cuts in Integer Linear Programming
por: Puigdemont, Pol, et al.
Publicado: (2024) -
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
por: Viel, Stefano, et al.
Publicado: (2025)