Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yibo, Zhao, Jiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fostering Intrinsic Motivation in Reinforcement Learning with Pretrained Foundation Models
von: Andres, Alain, et al.
Veröffentlicht: (2024)
von: Andres, Alain, et al.
Veröffentlicht: (2024)
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
von: Hugessen, Adriana, et al.
Veröffentlicht: (2024)
von: Hugessen, Adriana, et al.
Veröffentlicht: (2024)
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
von: Liu, Pai, et al.
Veröffentlicht: (2025)
von: Liu, Pai, et al.
Veröffentlicht: (2025)
Bootstrap Off-policy with World Model
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: a Short Survey
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
von: Colas, Cédric, et al.
Veröffentlicht: (2020)
Autotelic Reinforcement Learning: Exploring Intrinsic Motivations for Skill Acquisition in Open-Ended Environments
von: Srivastava, Prakhar, et al.
Veröffentlicht: (2025)
von: Srivastava, Prakhar, et al.
Veröffentlicht: (2025)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
Low Variance Off-policy Evaluation with State-based Importance Sampling
von: Bossens, David M., et al.
Veröffentlicht: (2022)
von: Bossens, David M., et al.
Veröffentlicht: (2022)
Image-Based Deep Reinforcement Learning with Intrinsically Motivated Stimuli: On the Execution of Complex Robotic Tasks
von: Valencia, David, et al.
Veröffentlicht: (2024)
von: Valencia, David, et al.
Veröffentlicht: (2024)
Improving Intrinsic Exploration by Creating Stationary Objectives
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2023)
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2023)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
von: Guiducci, Leonardo, et al.
Veröffentlicht: (2025)
von: Guiducci, Leonardo, et al.
Veröffentlicht: (2025)
MetaVIM: Meta Variationally Intrinsic Motivated Reinforcement Learning for Decentralized Traffic Signal Control
von: Zhu, Liwen, et al.
Veröffentlicht: (2021)
von: Zhu, Liwen, et al.
Veröffentlicht: (2021)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
von: Wang, Hanyong, et al.
Veröffentlicht: (2026)
Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction
von: He, Yiting, et al.
Veröffentlicht: (2025)
von: He, Yiting, et al.
Veröffentlicht: (2025)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
von: Chen, Kesheng, et al.
Veröffentlicht: (2025)
von: Chen, Kesheng, et al.
Veröffentlicht: (2025)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
Improving Offline-to-Online Reinforcement Learning with Q Conditioned State Entropy Exploration
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
OffSim: Offline Simulator for Model-based Offline Inverse Reinforcement Learning
von: Ahn, Woo-Jin, et al.
Veröffentlicht: (2025)
von: Ahn, Woo-Jin, et al.
Veröffentlicht: (2025)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
Causal Deepsets for Off-policy Evaluation under Spatial or Spatio-temporal Interferences
von: Dai, Runpeng, et al.
Veröffentlicht: (2024)
von: Dai, Runpeng, et al.
Veröffentlicht: (2024)
Adapting Critic Match Loss Landscape Visualization to Off-policy Reinforcement Learning
von: Liu, Jingyi, et al.
Veröffentlicht: (2026)
von: Liu, Jingyi, et al.
Veröffentlicht: (2026)
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
von: Wang, Likun, et al.
Veröffentlicht: (2025)
von: Wang, Likun, et al.
Veröffentlicht: (2025)
Primal-Dual Spectral Representation for Off-policy Evaluation
von: Hu, Yang, et al.
Veröffentlicht: (2024)
von: Hu, Yang, et al.
Veröffentlicht: (2024)
Rubric-based On-policy Distillation
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
von: Tu, Songjun, et al.
Veröffentlicht: (2024)
von: Tu, Songjun, et al.
Veröffentlicht: (2024)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
Neighboring State-based Exploration for Reinforcement Learning
von: Li, Yu-Teng, et al.
Veröffentlicht: (2022)
von: Li, Yu-Teng, et al.
Veröffentlicht: (2022)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
von: Romeo, Carlo, et al.
Veröffentlicht: (2026)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
von: Li, Yibo, et al.
Veröffentlicht: (2026)
von: Li, Yibo, et al.
Veröffentlicht: (2026)
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration
von: Yang, Yiqin, et al.
Veröffentlicht: (2026)
von: Yang, Yiqin, et al.
Veröffentlicht: (2026)
From Imitation to Exploration: End-to-end Autonomous Driving based on World Model
von: Li, Yueyuan, et al.
Veröffentlicht: (2024)
von: Li, Yueyuan, et al.
Veröffentlicht: (2024)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
von: Liang, Kun, et al.
Veröffentlicht: (2026)
von: Liang, Kun, et al.
Veröffentlicht: (2026)
A Pre-trained Data Deduplication Model based on Active Learning
von: Shi, Haochen, et al.
Veröffentlicht: (2023)
von: Shi, Haochen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Fostering Intrinsic Motivation in Reinforcement Learning with Pretrained Foundation Models
von: Andres, Alain, et al.
Veröffentlicht: (2024) -
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
von: Sun, Yan, et al.
Veröffentlicht: (2025) -
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
von: Liu, Zifan, et al.
Veröffentlicht: (2024) -
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
von: Hugessen, Adriana, et al.
Veröffentlicht: (2024) -
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
von: Liu, Pai, et al.
Veröffentlicht: (2025)