Learning Uncertainty-Aware Temporally-Extended Actions
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Joongkyu, Park, Seung Joon, Tang, Yunhao, Oh, Min-hwan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Combinatorial Reinforcement Learning with Preference Feedback
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
di: Lee, Joongkyu, et al.
Pubblicazione: (2024)
di: Lee, Joongkyu, et al.
Pubblicazione: (2024)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
di: Lee, Joongkyu, et al.
Pubblicazione: (2024)
di: Lee, Joongkyu, et al.
Pubblicazione: (2024)
Improved Online Confidence Bounds for Multinomial Logistic Bandits
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
di: Lee, Joongkyu, et al.
Pubblicazione: (2025)
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
di: Cho, Wooseong, et al.
Pubblicazione: (2024)
di: Cho, Wooseong, et al.
Pubblicazione: (2024)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
di: Yoon, Deokgyu, et al.
Pubblicazione: (2026)
di: Yoon, Deokgyu, et al.
Pubblicazione: (2026)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
di: Lee, Harin, et al.
Pubblicazione: (2025)
di: Lee, Harin, et al.
Pubblicazione: (2025)
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
di: Lee, Harin, et al.
Pubblicazione: (2026)
di: Lee, Harin, et al.
Pubblicazione: (2026)
Symmetry-Aware GFlowNets
di: Kim, Hohyun, et al.
Pubblicazione: (2025)
di: Kim, Hohyun, et al.
Pubblicazione: (2025)
Improved Regret of Linear Ensemble Sampling
di: Lee, Harin, et al.
Pubblicazione: (2024)
di: Lee, Harin, et al.
Pubblicazione: (2024)
Infrequent Exploration in Linear Bandits
di: Lee, Harin, et al.
Pubblicazione: (2025)
di: Lee, Harin, et al.
Pubblicazione: (2025)
Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation
di: Hwang, Taehyun, et al.
Pubblicazione: (2022)
di: Hwang, Taehyun, et al.
Pubblicazione: (2022)
Peng's Q($λ$) for Conservative Value Estimation in Offline Reinforcement Learning
di: Kim, Byeongchan, et al.
Pubblicazione: (2026)
di: Kim, Byeongchan, et al.
Pubblicazione: (2026)
Thompson Sampling for Multi-Objective Linear Contextual Bandit
di: Park, Somangchan, et al.
Pubblicazione: (2025)
di: Park, Somangchan, et al.
Pubblicazione: (2025)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
di: Kang, Hyungkyu, et al.
Pubblicazione: (2025)
di: Kang, Hyungkyu, et al.
Pubblicazione: (2025)
Blessings of Multiple Good Arms in Multi-Objective Linear Bandits
di: Ann, Heesang, et al.
Pubblicazione: (2026)
di: Ann, Heesang, et al.
Pubblicazione: (2026)
Optimal and Practical Batched Linear Bandit Algorithm
di: Yu, Sanghoon, et al.
Pubblicazione: (2025)
di: Yu, Sanghoon, et al.
Pubblicazione: (2025)
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
di: Yu, Sanghoon, et al.
Pubblicazione: (2026)
di: Yu, Sanghoon, et al.
Pubblicazione: (2026)
Lasso Bandit with Compatibility Condition on Optimal Arm
di: Lee, Harin, et al.
Pubblicazione: (2024)
di: Lee, Harin, et al.
Pubblicazione: (2024)
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
di: Kim, Chaiwon, et al.
Pubblicazione: (2025)
di: Kim, Chaiwon, et al.
Pubblicazione: (2025)
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
di: Kang, Hyungkyu, et al.
Pubblicazione: (2026)
di: Kang, Hyungkyu, et al.
Pubblicazione: (2026)
Queueing Matching Bandits with Preference Feedback
di: Kim, Jung-hun, et al.
Pubblicazione: (2024)
di: Kim, Jung-hun, et al.
Pubblicazione: (2024)
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
di: Kim, Seok-Jin, et al.
Pubblicazione: (2024)
di: Kim, Seok-Jin, et al.
Pubblicazione: (2024)
Exploration via Feature Perturbation in Contextual Bandits
di: Yi, Seouh-won, et al.
Pubblicazione: (2025)
di: Yi, Seouh-won, et al.
Pubblicazione: (2025)
ADAM Optimization with Adaptive Batch Selection
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2025)
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2025)
Stochastic Matching Bandits with Rare Optimization Updates
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
Dynamic Assortment Selection and Pricing with Censored Preference Feedback
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
Doubly Perturbed Task Free Continual Learning
di: Lee, Byung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Byung Hyun, et al.
Pubblicazione: (2023)
Follow-the-Perturbed-Leader with Fréchet-type Tail Distributions: Optimality in Adversarial Bandits and Best-of-Both-Worlds
di: Lee, Jongyeong, et al.
Pubblicazione: (2024)
di: Lee, Jongyeong, et al.
Pubblicazione: (2024)
Revisiting Follow-the-Perturbed-Leader with Unbounded Perturbations in Bandit Problems
di: Lee, Jongyeong, et al.
Pubblicazione: (2025)
di: Lee, Jongyeong, et al.
Pubblicazione: (2025)
Convergence of Muon with Newton-Schulz
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2026)
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2026)
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
di: Hwang, Taehyun, et al.
Pubblicazione: (2026)
di: Hwang, Taehyun, et al.
Pubblicazione: (2026)
Linear Bandits with Partially Observable Features
di: Kim, Wonyoung, et al.
Pubblicazione: (2025)
di: Kim, Wonyoung, et al.
Pubblicazione: (2025)
Oracle-Efficient Combinatorial Semi-Bandits
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
di: Kim, Jung-hun, et al.
Pubblicazione: (2025)
Benchmarking Uncertainty Disentanglement: Specialized Uncertainties for Specialized Tasks
di: Mucsányi, Bálint, et al.
Pubblicazione: (2024)
di: Mucsányi, Bálint, et al.
Pubblicazione: (2024)
Experimental Design for Semiparametric Bandits
di: Kim, Seok-Jin, et al.
Pubblicazione: (2025)
di: Kim, Seok-Jin, et al.
Pubblicazione: (2025)
Semantic-Aware Gaussian Process Calibration with Structured Layerwise Kernels for Deep Neural Networks
di: Lee, Kyung-hwan, et al.
Pubblicazione: (2025)
di: Lee, Kyung-hwan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Combinatorial Reinforcement Learning with Preference Feedback
di: Lee, Joongkyu, et al.
Pubblicazione: (2025) -
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
di: Lee, Joongkyu, et al.
Pubblicazione: (2024) -
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
di: Lee, Joongkyu, et al.
Pubblicazione: (2024) -
Improved Online Confidence Bounds for Multinomial Logistic Bandits
di: Lee, Joongkyu, et al.
Pubblicazione: (2025) -
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
di: Lee, Joongkyu, et al.
Pubblicazione: (2026)