Off-policy Reinforcement Learning with Model-based Exploration Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Likun, Zhang, Xiangteng, Wang, Yinuo, Zhan, Guojian, Wang, Wenxuan, Gao, Haoyu, Duan, Jingliang, Li, Shengbo Eben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bootstrap Off-policy with World Model
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
Conformal Symplectic Optimization for Stable Reinforcement Learning
von: Lyu, Yao, et al.
Veröffentlicht: (2024)
von: Lyu, Yao, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
Enhanced DACER Algorithm with High Diffusion Efficiency
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
von: Liu, Shiqi, et al.
Veröffentlicht: (2026)
von: Liu, Shiqi, et al.
Veröffentlicht: (2026)
Canonical Form of Datatic Description in Control Systems
von: Zhan, Guojian, et al.
Veröffentlicht: (2024)
von: Zhan, Guojian, et al.
Veröffentlicht: (2024)
Diffusion Actor-Critic with Entropy Regulator
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
Transferable Latent-to-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged Robots
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
von: Wang, Yibo, et al.
Veröffentlicht: (2024)
von: Wang, Yibo, et al.
Veröffentlicht: (2024)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning
von: Lu, Siyi, et al.
Veröffentlicht: (2024)
von: Lu, Siyi, et al.
Veröffentlicht: (2024)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
Zeroth-Order Actor-Critic: An Evolutionary Framework for Sequential Decision Problems
von: Lei, Yuheng, et al.
Veröffentlicht: (2022)
von: Lei, Yuheng, et al.
Veröffentlicht: (2022)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
von: Wang, Ruhan, et al.
Veröffentlicht: (2024)
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
An Explicit Discrete-Time Dynamic Vehicle Model with Assured Numerical Stability
von: Zhan, Guojian, et al.
Veröffentlicht: (2024)
von: Zhan, Guojian, et al.
Veröffentlicht: (2024)
Risk-Aware Vehicle Trajectory Prediction Under Safety-Critical Scenarios
von: Wang, Qingfan, et al.
Veröffentlicht: (2024)
von: Wang, Qingfan, et al.
Veröffentlicht: (2024)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
Minds on the Move: Decoding Trajectory Prediction in Autonomous Driving with Cognitive Insights
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
A Cognitive-Based Trajectory Prediction Approach for Autonomous Driving
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
von: Wang, Chen, et al.
Veröffentlicht: (2026)
von: Wang, Chen, et al.
Veröffentlicht: (2026)
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
Buffer Matters: Unleashing the Power of Off-Policy Reinforcement Learning in Large Language Model Reasoning
von: Wan, Xu, et al.
Veröffentlicht: (2026)
von: Wan, Xu, et al.
Veröffentlicht: (2026)
One Filters All: A Generalist Filter for State Estimation
von: Liu, Shiqi, et al.
Veröffentlicht: (2025)
von: Liu, Shiqi, et al.
Veröffentlicht: (2025)
Guardian: Decoupling Exploration from Safety in Reinforcement Learning
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
LocoMamba: Vision-Driven Locomotion via End-to-End Deep Reinforcement Learning with Mamba
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)
Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
von: Guan, Yang, et al.
Veröffentlicht: (2021)
von: Guan, Yang, et al.
Veröffentlicht: (2021)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
von: Li, Zeqiao, et al.
Veröffentlicht: (2026)
von: Li, Zeqiao, et al.
Veröffentlicht: (2026)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
von: Yang, Lu, et al.
Veröffentlicht: (2026)
von: Yang, Lu, et al.
Veröffentlicht: (2026)
Reinforcement Learning-based Sequential Route Recommendation for System-Optimal Traffic Assignment
von: Wang, Leizhen, et al.
Veröffentlicht: (2025)
von: Wang, Leizhen, et al.
Veröffentlicht: (2025)
Neighboring State-based Exploration for Reinforcement Learning
von: Li, Yu-Teng, et al.
Veröffentlicht: (2022)
von: Li, Yu-Teng, et al.
Veröffentlicht: (2022)
Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit
von: Huang, Tian, et al.
Veröffentlicht: (2023)
von: Huang, Tian, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Bootstrap Off-policy with World Model
von: Zhan, Guojian, et al.
Veröffentlicht: (2025) -
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
von: Zhan, Guojian, et al.
Veröffentlicht: (2025) -
Conformal Symplectic Optimization for Stable Reinforcement Learning
von: Lyu, Yao, et al.
Veröffentlicht: (2024) -
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025) -
Enhanced DACER Algorithm with High Diffusion Efficiency
von: Wang, Yinuo, et al.
Veröffentlicht: (2025)