Prior-Aligned Meta-RL: Thompson Sampling with Learned Priors and Guarantees in Finite-Horizon MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Runlin, Chen, Chixiang, Chen, Elynn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Prior Selection in Gaussian Process Bandits with Thompson Sampling
von: Sandberg, Jack, et al.
Veröffentlicht: (2025)
von: Sandberg, Jack, et al.
Veröffentlicht: (2025)
Fast Online Learning with Gaussian Prior-Driven Hierarchical Unimodal Thompson Sampling
von: Zhao, Tianchi, et al.
Veröffentlicht: (2026)
von: Zhao, Tianchi, et al.
Veröffentlicht: (2026)
Dynamic Prior Thompson Sampling for Cold-Start Exploration in Recommender Systems
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
von: Bayrooti, Jasmine, et al.
Veröffentlicht: (2025)
von: Bayrooti, Jasmine, et al.
Veröffentlicht: (2025)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Dual-Channel Tensor Neural Networks: Finite-Sample Theory and Conformal Structure Selection
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
High-Dimensional Tensor Discriminant Analysis: Low-Rank Discriminant Structure, Representation Synergy, and Theoretical Guarantees
von: Chen, Elynn, et al.
Veröffentlicht: (2025)
von: Chen, Elynn, et al.
Veröffentlicht: (2025)
One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
Addressing Finite-Horizon MDPs via Low-Rank Tensor Value Approximation
von: Rozada, Sergio, et al.
Veröffentlicht: (2025)
von: Rozada, Sergio, et al.
Veröffentlicht: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
PFN-TS: Thompson Sampling for Contextual Bandits via Prior-Data Fitted Networks
von: Tan, Yan Shuo, et al.
Veröffentlicht: (2026)
von: Tan, Yan Shuo, et al.
Veröffentlicht: (2026)
Amortising Inference and Meta-Learning Priors in Neural Networks
von: Rochussen, Tommy, et al.
Veröffentlicht: (2026)
von: Rochussen, Tommy, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning
von: Chai, Jinhang, et al.
Veröffentlicht: (2025)
von: Chai, Jinhang, et al.
Veröffentlicht: (2025)
Data-Driven Knowledge Transfer in Batch $Q^*$ Learning
von: Chen, Elynn, et al.
Veröffentlicht: (2024)
von: Chen, Elynn, et al.
Veröffentlicht: (2024)
Transfer Learning for Contextual Joint Assortment-Pricing under Cross-Market Heterogeneity
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
Language-Induced Priors for Domain Adaptation
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
Thompson Sampling for Repeated Newsvendor
von: Chen, Li, et al.
Veröffentlicht: (2025)
von: Chen, Li, et al.
Veröffentlicht: (2025)
Transition Transfer $Q$-Learning for Composite Markov Decision Processes
von: Chai, Jinhang, et al.
Veröffentlicht: (2025)
von: Chai, Jinhang, et al.
Veröffentlicht: (2025)
Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift
von: Li, Bochao, et al.
Veröffentlicht: (2026)
von: Li, Bochao, et al.
Veröffentlicht: (2026)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture Priors
von: Sefidgaran, Milad, et al.
Veröffentlicht: (2025)
von: Sefidgaran, Milad, et al.
Veröffentlicht: (2025)
Aligning Latent Spaces with Flow Priors
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
von: Mutti, Mirco, et al.
Veröffentlicht: (2023)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
CAESAR: Enhancing Federated RL in Heterogeneous MDPs through Convergence-Aware Sampling with Screening
von: Mak, Hei Yi, et al.
Veröffentlicht: (2024)
von: Mak, Hei Yi, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Control with Probabilistic Stability Guarantee: A Finite-Sample Approach
von: Han, Minghao, et al.
Veröffentlicht: (2026)
von: Han, Minghao, et al.
Veröffentlicht: (2026)
Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
Online Posterior Sampling with a Diffusion Prior
von: Kveton, Branislav, et al.
Veröffentlicht: (2024)
von: Kveton, Branislav, et al.
Veröffentlicht: (2024)
Regenerative Particle Thompson Sampling
von: Zhou, Zeyu, et al.
Veröffentlicht: (2022)
von: Zhou, Zeyu, et al.
Veröffentlicht: (2022)
ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning
von: Zhu, Tong, et al.
Veröffentlicht: (2026)
von: Zhu, Tong, et al.
Veröffentlicht: (2026)
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
von: Lazzaro, Joseph, et al.
Veröffentlicht: (2026)
von: Lazzaro, Joseph, et al.
Veröffentlicht: (2026)
Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding
von: Xu, Jiamin, et al.
Veröffentlicht: (2026)
von: Xu, Jiamin, et al.
Veröffentlicht: (2026)
Solving robust MDPs as a sequence of static RL problems
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
von: Vora, Manav, et al.
Veröffentlicht: (2024)
von: Vora, Manav, et al.
Veröffentlicht: (2024)
Online Learning of Decision Trees with Thompson Sampling
von: Chaouki, Ayman, et al.
Veröffentlicht: (2024)
von: Chaouki, Ayman, et al.
Veröffentlicht: (2024)
Divide-and-Conquer Posterior Sampling for Denoising Diffusion Priors
von: Janati, Yazid, et al.
Veröffentlicht: (2024)
von: Janati, Yazid, et al.
Veröffentlicht: (2024)
Tensor-Fused Multi-View Graph Contrastive Learning
von: Wu, Yujia, et al.
Veröffentlicht: (2024)
von: Wu, Yujia, et al.
Veröffentlicht: (2024)
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making
von: Chen, Ruoyu, et al.
Veröffentlicht: (2026)
von: Chen, Ruoyu, et al.
Veröffentlicht: (2026)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Adaptive Prior Selection in Gaussian Process Bandits with Thompson Sampling
von: Sandberg, Jack, et al.
Veröffentlicht: (2025) -
Fast Online Learning with Gaussian Prior-Driven Hierarchical Unimodal Thompson Sampling
von: Zhao, Tianchi, et al.
Veröffentlicht: (2026) -
Dynamic Prior Thompson Sampling for Cold-Start Exploration in Recommender Systems
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026) -
No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
von: Bayrooti, Jasmine, et al.
Veröffentlicht: (2025) -
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
von: Chen, Xin, et al.
Veröffentlicht: (2024)