Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boudart, Pierre, Gaillard, Pierre, Rudi, Alessandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
von: Boudart, Pierre, et al.
Veröffentlicht: (2025)
von: Boudart, Pierre, et al.
Veröffentlicht: (2025)
Structured Prediction in Online Learning
von: Boudart, Pierre, et al.
Veröffentlicht: (2024)
von: Boudart, Pierre, et al.
Veröffentlicht: (2024)
Minimax-optimal and Locally-adaptive Online Nonparametric Regression
von: Liautaud, Paul, et al.
Veröffentlicht: (2024)
von: Liautaud, Paul, et al.
Veröffentlicht: (2024)
Minimax Adaptive Online Nonparametric Regression over Besov Spaces
von: Liautaud, Paul, et al.
Veröffentlicht: (2025)
von: Liautaud, Paul, et al.
Veröffentlicht: (2025)
High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch
von: Liautaud, Paul, et al.
Veröffentlicht: (2026)
von: Liautaud, Paul, et al.
Veröffentlicht: (2026)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
von: Yan, Fanqi, et al.
Veröffentlicht: (2026)
von: Yan, Fanqi, et al.
Veröffentlicht: (2026)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
Near-Optimal Learning and Planning in Separated Latent MDPs
von: Chen, Fan, et al.
Veröffentlicht: (2024)
von: Chen, Fan, et al.
Veröffentlicht: (2024)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
Variance-Aware Estimation of Kernel Mean Embedding
von: Wolfer, Geoffrey, et al.
Veröffentlicht: (2022)
von: Wolfer, Geoffrey, et al.
Veröffentlicht: (2022)
Precise Asymptotics and Refined Regret of Variance-Aware UCB
von: Fan, Yingying, et al.
Veröffentlicht: (2024)
von: Fan, Yingying, et al.
Veröffentlicht: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
von: Zhou, Julien, et al.
Veröffentlicht: (2024)
Minimax Optimal Simple Regret in Two-Armed Best-Arm Identification
von: Kato, Masahiro
Veröffentlicht: (2024)
von: Kato, Masahiro
Veröffentlicht: (2024)
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
von: Réveillard, William, et al.
Veröffentlicht: (2025)
von: Réveillard, William, et al.
Veröffentlicht: (2025)
Minimax Regret Learning for Data with Heterogeneous Subgroups
von: Mo, Weibin, et al.
Veröffentlicht: (2024)
von: Mo, Weibin, et al.
Veröffentlicht: (2024)
Minimax Optimal Fair Classification with Bounded Demographic Disparity
von: Zeng, Xianli, et al.
Veröffentlicht: (2024)
von: Zeng, Xianli, et al.
Veröffentlicht: (2024)
Optimal rates for density and mode estimation with expand-and-sparsify representations
von: Sinha, Kaushik, et al.
Veröffentlicht: (2026)
von: Sinha, Kaushik, et al.
Veröffentlicht: (2026)
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
von: Roy, Saptarshi, et al.
Veröffentlicht: (2026)
von: Roy, Saptarshi, et al.
Veröffentlicht: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
Minimax Optimality of Score-based Diffusion Models: Beyond the Density Lower Bound Assumptions
von: Zhang, Kaihong, et al.
Veröffentlicht: (2024)
von: Zhang, Kaihong, et al.
Veröffentlicht: (2024)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
von: Hao, Sai, et al.
Veröffentlicht: (2026)
von: Hao, Sai, et al.
Veröffentlicht: (2026)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
von: Yu, Hao
Veröffentlicht: (2025)
von: Yu, Hao
Veröffentlicht: (2025)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
MetaCURL: Non-stationary Concave Utility Reinforcement Learning
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2024)
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2024)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
von: He, Jiafan, et al.
Veröffentlicht: (2025)
von: He, Jiafan, et al.
Veröffentlicht: (2025)
Statistical Inference for Optimal Transport Maps: Recent Advances and Perspectives
von: Balakrishnan, Sivaraman, et al.
Veröffentlicht: (2025)
von: Balakrishnan, Sivaraman, et al.
Veröffentlicht: (2025)
Inference with the Upper Confidence Bound Algorithm
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
von: Khamaru, Koulik, et al.
Veröffentlicht: (2024)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
Bias-Aware Conformal Prediction for Metric-Based Imaging Pipelines
von: Cheung, Matt Y., et al.
Veröffentlicht: (2024)
von: Cheung, Matt Y., et al.
Veröffentlicht: (2024)
No-Regret Reinforcement Learning in Smooth MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
Convergence of Statistical Estimators via Mutual Information Bounds
von: Khribch, El Mahdi, et al.
Veröffentlicht: (2024)
von: Khribch, El Mahdi, et al.
Veröffentlicht: (2024)
RMLR: Extending Multinomial Logistic Regression into General Geometries
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
von: Chen, Ziheng, et al.
Veröffentlicht: (2024)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
von: Praharaj, Samya, et al.
Veröffentlicht: (2025)
Minimax Optimality of the Probability Flow ODE for Diffusion Models
von: Cai, Changxiao, et al.
Veröffentlicht: (2025)
von: Cai, Changxiao, et al.
Veröffentlicht: (2025)
Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
von: Chen, Fan, et al.
Veröffentlicht: (2026)
von: Chen, Fan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
von: Boudart, Pierre, et al.
Veröffentlicht: (2025) -
Structured Prediction in Online Learning
von: Boudart, Pierre, et al.
Veröffentlicht: (2024) -
Minimax-optimal and Locally-adaptive Online Nonparametric Regression
von: Liautaud, Paul, et al.
Veröffentlicht: (2024) -
Minimax Adaptive Online Nonparametric Regression over Besov Spaces
von: Liautaud, Paul, et al.
Veröffentlicht: (2025) -
High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch
von: Liautaud, Paul, et al.
Veröffentlicht: (2026)