Improving Offline RL by Blending Heuristics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Geng, Sinong, Pacchiano, Aldo, Kolobov, Andrey, Cheng, Ching-An |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)
von: Pacchiano, Aldo
Veröffentlicht: (2024)
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
von: Hu, Xinyi, et al.
Veröffentlicht: (2025)
von: Hu, Xinyi, et al.
Veröffentlicht: (2025)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
In-Context Learning for Pure Exploration in Continuous Spaces
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
Bayesian Online Model Selection
von: Afshar, Aida, et al.
Veröffentlicht: (2026)
von: Afshar, Aida, et al.
Veröffentlicht: (2026)
Pure Exploration with Feedback Graphs
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Contextual Bandits with Stage-wise Constraints
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control
von: Zheng, Ruijie, et al.
Veröffentlicht: (2024)
von: Zheng, Ruijie, et al.
Veröffentlicht: (2024)
Data-Driven Online Model Selection With Regret Guarantees
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
State-free Reinforcement Learning
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
In-Context Learning for Pure Exploration
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
On the Hardness of Bandit Learning
von: Brukhim, Nataly, et al.
Veröffentlicht: (2025)
von: Brukhim, Nataly, et al.
Veröffentlicht: (2025)
Language Model Personalization via Reward Factorization
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
Multiple-policy Evaluation via Density Estimation
von: Chen, Yilei, et al.
Veröffentlicht: (2024)
von: Chen, Yilei, et al.
Veröffentlicht: (2024)
Experiment Planning with Function Approximation
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
Provable Interactive Learning with Hindsight Instruction Feedback
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
A Theoretical Framework for Partially Observed Reward-States in RLHF
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
von: Luo, Yu, et al.
Veröffentlicht: (2024)
von: Luo, Yu, et al.
Veröffentlicht: (2024)
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
von: Landers, Matthew, et al.
Veröffentlicht: (2026)
von: Landers, Matthew, et al.
Veröffentlicht: (2026)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
When Are RL Hyperparameters Benign? A Study in Offline Goal-Conditioned RL
von: Töpperwien, Jan Malte, et al.
Veröffentlicht: (2026)
von: Töpperwien, Jan Malte, et al.
Veröffentlicht: (2026)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
von: Wang, Qi, et al.
Veröffentlicht: (2023)
von: Wang, Qi, et al.
Veröffentlicht: (2023)
Selective Uncertainty Propagation in Offline RL
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
von: Krishnamurthy, Sanath Kumar, et al.
Veröffentlicht: (2023)
Decoupled Prioritized Resampling for Offline RL
von: Yue, Yang, et al.
Veröffentlicht: (2023)
von: Yue, Yang, et al.
Veröffentlicht: (2023)
Augmenting Offline RL with Unlabeled Data
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
von: Gupta, Aaryan, et al.
Veröffentlicht: (2025)
von: Gupta, Aaryan, et al.
Veröffentlicht: (2025)
Offline RL via Feature-Occupancy Gradient Ascent
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
Reinformer: Max-Return Sequence Modeling for Offline RL
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
Latent Policy Steering through One-Step Flow Policies
von: Im, Hokyun, et al.
Veröffentlicht: (2026)
von: Im, Hokyun, et al.
Veröffentlicht: (2026)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
von: Su, Jianhai, et al.
Veröffentlicht: (2025)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
von: Fan, Ying, et al.
Veröffentlicht: (2024)
von: Fan, Ying, et al.
Veröffentlicht: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
von: Mark, Max Sobol, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024) -
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
von: Afshar, Aida, et al.
Veröffentlicht: (2025) -
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
von: Hu, Xinyi, et al.
Veröffentlicht: (2025) -
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025) -
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
von: Afshar, Aida, et al.
Veröffentlicht: (2024)