Scaling Offline RL via Efficient and Expressive Shortcut Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Espinosa-Dice, Nicolas, Zhang, Yiyi, Chen, Yiding, Guo, Bradley, Oertell, Owen, Swamy, Gokul, Brantley, Kiante, Sun, Wen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915311560687616
author Espinosa-Dice, Nicolas
Zhang, Yiyi
Chen, Yiding
Guo, Bradley
Oertell, Owen
Swamy, Gokul
Brantley, Kiante
Sun, Wen
author_facet Espinosa-Dice, Nicolas
Zhang, Yiyi
Chen, Yiding
Guo, Bradley
Oertell, Owen
Swamy, Gokul
Brantley, Kiante
Sun, Wen
contents Diffusion and flow models have emerged as powerful generative approaches capable of modeling diverse and multimodal behavior. However, applying these models to offline reinforcement learning (RL) remains challenging due to the iterative nature of their noise sampling processes, making policy optimization difficult. In this paper, we introduce Scalable Offline Reinforcement Learning (SORL), a new offline RL algorithm that leverages shortcut models - a novel class of generative models - to scale both training and inference. SORL's policy can capture complex data distributions and can be trained simply and efficiently in a one-stage training procedure. At test time, SORL introduces both sequential and parallel inference scaling by using the learned Q-function as a verifier. We demonstrate that SORL achieves strong performance across a range of offline RL tasks and exhibits positive scaling behavior with increased test-time compute. We release the code at nico-espinosadice.github.io/projects/sorl.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Offline RL via Efficient and Expressive Shortcut Models
Espinosa-Dice, Nicolas
Zhang, Yiyi
Chen, Yiding
Guo, Bradley
Oertell, Owen
Swamy, Gokul
Brantley, Kiante
Sun, Wen
Machine Learning
Artificial Intelligence
I.2.6
Diffusion and flow models have emerged as powerful generative approaches capable of modeling diverse and multimodal behavior. However, applying these models to offline reinforcement learning (RL) remains challenging due to the iterative nature of their noise sampling processes, making policy optimization difficult. In this paper, we introduce Scalable Offline Reinforcement Learning (SORL), a new offline RL algorithm that leverages shortcut models - a novel class of generative models - to scale both training and inference. SORL's policy can capture complex data distributions and can be trained simply and efficiently in a one-stage training procedure. At test time, SORL introduces both sequential and parallel inference scaling by using the learned Q-function as a verifier. We demonstrate that SORL achieves strong performance across a range of offline RL tasks and exhibits positive scaling behavior with increased test-time compute. We release the code at nico-espinosadice.github.io/projects/sorl.
title Scaling Offline RL via Efficient and Expressive Shortcut Models
topic Machine Learning
Artificial Intelligence
I.2.6
url https://arxiv.org/abs/2505.22866