Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Bozkurt, Alper Kamil, Xu, Xiaoan, Zhang, Shangtong, Pajic, Miroslav, Motai, Yuichi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning
by: Bozkurt, Alper Kamil, et al.
Published: (2019)
by: Bozkurt, Alper Kamil, et al.
Published: (2019)
Learning Optimal Strategies for Temporal Tasks in Stochastic Games
by: Bozkurt, Alper Kamil, et al.
Published: (2021)
by: Bozkurt, Alper Kamil, et al.
Published: (2021)
Neuro-Logic Lifelong Learning
by: He, Bowen, et al.
Published: (2025)
by: He, Bowen, et al.
Published: (2025)
Safe In-Context Reinforcement Learning
by: Moeini, Amir, et al.
Published: (2025)
by: Moeini, Amir, et al.
Published: (2025)
On the Uniqueness of Solution for the Bellman Equation of LTL Objectives
by: Xuan, Zetong, et al.
Published: (2024)
by: Xuan, Zetong, et al.
Published: (2024)
Model-Free Learning of Safe yet Effective Controllers
by: Bozkurt, Alper Kamil, et al.
Published: (2021)
by: Bozkurt, Alper Kamil, et al.
Published: (2021)
Secure Planning Against Stealthy Attacks via Model-Free Reinforcement Learning
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
Model-Free Reinforcement Learning for Stochastic Games with Linear Temporal Logic Objectives
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
by: Bozkurt, Alper Kamil, et al.
Published: (2020)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
by: Xiao, Wei, et al.
Published: (2025)
by: Xiao, Wei, et al.
Published: (2025)
Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning
by: Song, Chihyeon, et al.
Published: (2025)
by: Song, Chihyeon, et al.
Published: (2025)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
by: Madhow, Sunil, et al.
Published: (2023)
by: Madhow, Sunil, et al.
Published: (2023)
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning
by: Chemingui, Yassine, et al.
Published: (2024)
by: Chemingui, Yassine, et al.
Published: (2024)
Off-Policy Selection for Initiating Human-Centric Experimental Design
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
by: Yang, Letian, et al.
Published: (2026)
by: Yang, Letian, et al.
Published: (2026)
Selective Reincarnation: Offline-to-Online Multi-Agent Reinforcement Learning
by: Formanek, Claude, et al.
Published: (2023)
by: Formanek, Claude, et al.
Published: (2023)
Safe Offline Reinforcement Learning with Real-Time Budget Constraints
by: Lin, Qian, et al.
Published: (2023)
by: Lin, Qian, et al.
Published: (2023)
SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning
by: Xu, Zhongling, et al.
Published: (2026)
by: Xu, Zhongling, et al.
Published: (2026)
Offline Reinforcement Learning with Behavioral Supervisor Tuning
by: Srinivasan, Padmanaba, et al.
Published: (2024)
by: Srinivasan, Padmanaba, et al.
Published: (2024)
Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation
by: Huang, Xiao, et al.
Published: (2025)
by: Huang, Xiao, et al.
Published: (2025)
Counterfactual Explanations for Continuous Action Reinforcement Learning
by: Dong, Shuyang, et al.
Published: (2025)
by: Dong, Shuyang, et al.
Published: (2025)
Online Optimization for Offline Safe Reinforcement Learning
by: Chemingui, Yassine, et al.
Published: (2025)
by: Chemingui, Yassine, et al.
Published: (2025)
The Three Regimes of Offline-to-Online Reinforcement Learning
by: Li, Lu, et al.
Published: (2025)
by: Li, Lu, et al.
Published: (2025)
SAMG: Offline-to-Online Reinforcement Learning via State-Action-Conditional Offline Model Guidance
by: Zhang, Liyu, et al.
Published: (2024)
by: Zhang, Liyu, et al.
Published: (2024)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
by: Nakamoto, Mitsuhiko, et al.
Published: (2023)
Offline Reinforcement Learning with Generative Trajectory Policies
by: Feng, Xinsong, et al.
Published: (2025)
by: Feng, Xinsong, et al.
Published: (2025)
Policy Expansion for Bridging Offline-to-Online Reinforcement Learning
by: Zhang, Haichao, et al.
Published: (2023)
by: Zhang, Haichao, et al.
Published: (2023)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
by: Ma, Lu, et al.
Published: (2025)
by: Ma, Lu, et al.
Published: (2025)
TSN-Affinity: Similarity-Driven Parameter Reuse for Continual Offline Reinforcement Learning
by: Żurek, Dominik, et al.
Published: (2026)
by: Żurek, Dominik, et al.
Published: (2026)
Discrete Flow Matching for Offline-to-Online Reinforcement Learning
by: Khan, Fairoz Nower, et al.
Published: (2026)
by: Khan, Fairoz Nower, et al.
Published: (2026)
Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2026)
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2026)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
by: Hu, Hao, et al.
Published: (2025)
by: Hu, Hao, et al.
Published: (2025)
Policy-regularized Offline Multi-objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024)
by: Lin, Qian, et al.
Published: (2024)
Evaluation-Time Policy Switching for Offline Reinforcement Learning
by: Neggatu, Natinael Solomon, et al.
Published: (2025)
by: Neggatu, Natinael Solomon, et al.
Published: (2025)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
by: He, Longxiang, et al.
Published: (2025)
by: He, Longxiang, et al.
Published: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
Similar Items
-
Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning
by: Bozkurt, Alper Kamil, et al.
Published: (2019) -
Learning Optimal Strategies for Temporal Tasks in Stochastic Games
by: Bozkurt, Alper Kamil, et al.
Published: (2021) -
Neuro-Logic Lifelong Learning
by: He, Bowen, et al.
Published: (2025) -
Safe In-Context Reinforcement Learning
by: Moeini, Amir, et al.
Published: (2025) -
On the Uniqueness of Solution for the Bellman Equation of LTL Objectives
by: Xuan, Zetong, et al.
Published: (2024)