Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Wonyoung, Oh, Min-Hwan, Iyengar, Garud, Zeevi, Assaf |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning the Pareto Front Using Bootstrapped Observation Samples
by: Kim, Wonyoung, et al.
Published: (2023)
by: Kim, Wonyoung, et al.
Published: (2023)
Linear Bandits with Partially Observable Features
by: Kim, Wonyoung, et al.
Published: (2025)
by: Kim, Wonyoung, et al.
Published: (2025)
The Cost of Learning under Multiple Change Points
by: Gafni, Tomer, et al.
Published: (2026)
by: Gafni, Tomer, et al.
Published: (2026)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation
by: Hwang, Taehyun, et al.
Published: (2022)
by: Hwang, Taehyun, et al.
Published: (2022)
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
by: Lee, Joongkyu, et al.
Published: (2026)
by: Lee, Joongkyu, et al.
Published: (2026)
Model-Free Approximate Bayesian Learning for Large-Scale Conversion Funnel Optimization
by: Iyengar, Garud, et al.
Published: (2024)
by: Iyengar, Garud, et al.
Published: (2024)
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
by: Cho, Wooseong, et al.
Published: (2024)
by: Cho, Wooseong, et al.
Published: (2024)
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
by: Hwang, Taehyun, et al.
Published: (2026)
by: Hwang, Taehyun, et al.
Published: (2026)
A Tractable Online Learning Algorithm for the Multinomial Logit Contextual Bandit
by: Agrawal, Priyank, et al.
Published: (2020)
by: Agrawal, Priyank, et al.
Published: (2020)
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Adaptive Querying with AI Persona Priors
by: Wang, Kaizheng, et al.
Published: (2026)
by: Wang, Kaizheng, et al.
Published: (2026)
Bayesian Design Principles for Frequentist Sequential Learning
by: Xu, Yunbei, et al.
Published: (2023)
by: Xu, Yunbei, et al.
Published: (2023)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Adaptive Data Augmentation for Thompson Sampling
by: Kim, Wonyoung
Published: (2025)
by: Kim, Wonyoung
Published: (2025)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Upper Counterfactual Confidence Bounds: a New Optimism Principle for Contextual Bandits
by: Xu, Yunbei, et al.
Published: (2020)
by: Xu, Yunbei, et al.
Published: (2020)
Optimizer's Information Criterion: Dissecting and Correcting Bias in Data-Driven Optimization
by: Iyengar, Garud, et al.
Published: (2023)
by: Iyengar, Garud, et al.
Published: (2023)
A Broader View of Thompson Sampling
by: Qu, Yanlin, et al.
Published: (2025)
by: Qu, Yanlin, et al.
Published: (2025)
Improved Online Confidence Bounds for Multinomial Logistic Bandits
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Learning Multinomial Logits in $O(n \log n)$ time
by: Chierichetti, Flavio, et al.
Published: (2026)
by: Chierichetti, Flavio, et al.
Published: (2026)
Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
Model-Free Assessment of Simulator Fidelity via Quantile Curves
by: Iyengar, Garud, et al.
Published: (2025)
by: Iyengar, Garud, et al.
Published: (2025)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
by: Boudart, Pierre, et al.
Published: (2026)
by: Boudart, Pierre, et al.
Published: (2026)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Adaptive Resolving Methods for Reinforcement Learning with Function Approximations
by: Jiang, Jiashuo, et al.
Published: (2025)
by: Jiang, Jiashuo, et al.
Published: (2025)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
by: Zhao, Heyang, et al.
Published: (2023)
by: Zhao, Heyang, et al.
Published: (2023)
Optimal and Practical Batched Linear Bandit Algorithm
by: Yu, Sanghoon, et al.
Published: (2025)
by: Yu, Sanghoon, et al.
Published: (2025)
Peng's Q($λ$) for Conservative Value Estimation in Offline Reinforcement Learning
by: Kim, Byeongchan, et al.
Published: (2026)
by: Kim, Byeongchan, et al.
Published: (2026)
Achieving Limited Adaptivity for Multinomial Logistic Bandits
by: Midigeshi, Sukruta Prakash, et al.
Published: (2025)
by: Midigeshi, Sukruta Prakash, et al.
Published: (2025)
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
by: Yu, Sanghoon, et al.
Published: (2026)
by: Yu, Sanghoon, et al.
Published: (2026)
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
by: Kang, Hyungkyu, et al.
Published: (2026)
by: Kang, Hyungkyu, et al.
Published: (2026)
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
FIRAL: An Active Learning Algorithm for Multinomial Logistic Regression
by: Chen, Youguang, et al.
Published: (2024)
by: Chen, Youguang, et al.
Published: (2024)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
Combinatorial Reinforcement Learning with Preference Feedback
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
ADAM Optimization with Adaptive Batch Selection
by: Kim, Gyu Yeol, et al.
Published: (2025)
by: Kim, Gyu Yeol, et al.
Published: (2025)
Nonstationary Reinforcement Learning with Linear Function Approximation
by: Zhou, Huozhi, et al.
Published: (2020)
by: Zhou, Huozhi, et al.
Published: (2020)
Replicable Reinforcement Learning with Linear Function Approximation
by: Eaton, Eric, et al.
Published: (2025)
by: Eaton, Eric, et al.
Published: (2025)
A Temporally Correlated Latent Exploration for Reinforcement Learning
by: Oh, SuMin, et al.
Published: (2024)
by: Oh, SuMin, et al.
Published: (2024)
Similar Items
-
Learning the Pareto Front Using Bootstrapped Observation Samples
by: Kim, Wonyoung, et al.
Published: (2023) -
Linear Bandits with Partially Observable Features
by: Kim, Wonyoung, et al.
Published: (2025) -
The Cost of Learning under Multiple Change Points
by: Gafni, Tomer, et al.
Published: (2026) -
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024) -
Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation
by: Hwang, Taehyun, et al.
Published: (2022)