SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Yao-Hui, Wang, Zeyu, Li, Xin, Pang, Wei, Yuan, Yingfang, Chen, Zhengkun, Zhang, Boya, Islam, Riashat, Lamb, Alex, Zhang, Yonggang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915990329098240
author Li, Yao-Hui
Wang, Zeyu
Li, Xin
Pang, Wei
Yuan, Yingfang
Chen, Zhengkun
Zhang, Boya
Islam, Riashat
Lamb, Alex
Zhang, Yonggang
author_facet Li, Yao-Hui
Wang, Zeyu
Li, Xin
Pang, Wei
Yuan, Yingfang
Chen, Zhengkun
Zhang, Boya
Islam, Riashat
Lamb, Alex
Zhang, Yonggang
contents Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose Shaping Landscapes with Optimistic Potential Estimates (SLOPE), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on 30+ tasks across 5 benchmarks and real-world robotic deployments, demonstrate that SLOPE consistently outperforms leading baselines in fully sparse, semi-sparse, and dense rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03201
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
Li, Yao-Hui
Wang, Zeyu
Li, Xin
Pang, Wei
Yuan, Yingfang
Chen, Zhengkun
Zhang, Boya
Islam, Riashat
Lamb, Alex
Zhang, Yonggang
Machine Learning
Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose Shaping Landscapes with Optimistic Potential Estimates (SLOPE), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on 30+ tasks across 5 benchmarks and real-world robotic deployments, demonstrate that SLOPE consistently outperforms leading baselines in fully sparse, semi-sparse, and dense rewards.
title SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2602.03201