Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Taehyun, Oh, Min-hwan |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
by: Cho, Wooseong, et al.
Published: (2024)
by: Cho, Wooseong, et al.
Published: (2024)
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
by: Hwang, Taehyun, et al.
Published: (2026)
by: Hwang, Taehyun, et al.
Published: (2026)
Improved Online Confidence Bounds for Multinomial Logistic Bandits
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
by: Lee, Joongkyu, et al.
Published: (2026)
by: Lee, Joongkyu, et al.
Published: (2026)
Lasso Bandit with Compatibility Condition on Optimal Arm
by: Lee, Harin, et al.
Published: (2024)
by: Lee, Harin, et al.
Published: (2024)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation
by: Kim, Wonyoung, et al.
Published: (2026)
by: Kim, Wonyoung, et al.
Published: (2026)
Combinatorial Reinforcement Learning with Preference Feedback
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
by: Lee, Harin, et al.
Published: (2026)
by: Lee, Harin, et al.
Published: (2026)
Peng's Q($λ$) for Conservative Value Estimation in Offline Reinforcement Learning
by: Kim, Byeongchan, et al.
Published: (2026)
by: Kim, Byeongchan, et al.
Published: (2026)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
by: Kang, Hyungkyu, et al.
Published: (2025)
by: Kang, Hyungkyu, et al.
Published: (2025)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
by: Kang, Hyungkyu, et al.
Published: (2026)
by: Kang, Hyungkyu, et al.
Published: (2026)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
by: Pham, Tuan Minh, et al.
Published: (2026)
by: Pham, Tuan Minh, et al.
Published: (2026)
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Blessings of Multiple Good Arms in Multi-Objective Linear Bandits
by: Ann, Heesang, et al.
Published: (2026)
by: Ann, Heesang, et al.
Published: (2026)
Infrequent Exploration in Linear Bandits
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Optimal and Practical Batched Linear Bandit Algorithm
by: Yu, Sanghoon, et al.
Published: (2025)
by: Yu, Sanghoon, et al.
Published: (2025)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
by: Lee, Joongkyu, et al.
Published: (2026)
by: Lee, Joongkyu, et al.
Published: (2026)
Improved Regret of Linear Ensemble Sampling
by: Lee, Harin, et al.
Published: (2024)
by: Lee, Harin, et al.
Published: (2024)
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
by: Yu, Sanghoon, et al.
Published: (2026)
by: Yu, Sanghoon, et al.
Published: (2026)
Achieving Limited Adaptivity for Multinomial Logistic Bandits
by: Midigeshi, Sukruta Prakash, et al.
Published: (2025)
by: Midigeshi, Sukruta Prakash, et al.
Published: (2025)
FIRAL: An Active Learning Algorithm for Multinomial Logistic Regression
by: Chen, Youguang, et al.
Published: (2024)
by: Chen, Youguang, et al.
Published: (2024)
Riemannian Multinomial Logistics Regression for SPD Neural Networks
by: Chen, Ziheng, et al.
Published: (2023)
by: Chen, Ziheng, et al.
Published: (2023)
Exploration via Feature Perturbation in Contextual Bandits
by: Yi, Seouh-won, et al.
Published: (2025)
by: Yi, Seouh-won, et al.
Published: (2025)
ADAM Optimization with Adaptive Batch Selection
by: Kim, Gyu Yeol, et al.
Published: (2025)
by: Kim, Gyu Yeol, et al.
Published: (2025)
Queueing Matching Bandits with Preference Feedback
by: Kim, Jung-hun, et al.
Published: (2024)
by: Kim, Jung-hun, et al.
Published: (2024)
Stochastic Matching Bandits with Rare Optimization Updates
by: Kim, Jung-hun, et al.
Published: (2025)
by: Kim, Jung-hun, et al.
Published: (2025)
Dynamic Assortment Selection and Pricing with Censored Preference Feedback
by: Kim, Jung-hun, et al.
Published: (2025)
by: Kim, Jung-hun, et al.
Published: (2025)
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
by: Kim, Seok-Jin, et al.
Published: (2024)
by: Kim, Seok-Jin, et al.
Published: (2024)
RMLR: Extending Multinomial Logistic Regression into General Geometries
by: Chen, Ziheng, et al.
Published: (2024)
by: Chen, Ziheng, et al.
Published: (2024)
Convergence of Muon with Newton-Schulz
by: Kim, Gyu Yeol, et al.
Published: (2026)
by: Kim, Gyu Yeol, et al.
Published: (2026)
Scalable Inference for Bayesian Multinomial Logistic-Normal Dynamic Linear Models
by: Saxena, Manan, et al.
Published: (2024)
by: Saxena, Manan, et al.
Published: (2024)
Thompson Sampling for Multi-Objective Linear Contextual Bandit
by: Park, Somangchan, et al.
Published: (2025)
by: Park, Somangchan, et al.
Published: (2025)
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
by: Kim, Chaiwon, et al.
Published: (2025)
by: Kim, Chaiwon, et al.
Published: (2025)
Symmetry-Aware GFlowNets
by: Kim, Hohyun, et al.
Published: (2025)
by: Kim, Hohyun, et al.
Published: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Similar Items
-
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
by: Cho, Wooseong, et al.
Published: (2024) -
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
by: Hwang, Taehyun, et al.
Published: (2026) -
Improved Online Confidence Bounds for Multinomial Logistic Bandits
by: Lee, Joongkyu, et al.
Published: (2025) -
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
by: Lee, Joongkyu, et al.
Published: (2024) -
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
by: Lee, Joongkyu, et al.
Published: (2026)