Drift Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Houssaini, Anas, Danesh, Mohamad H., Abyaneh, Amin, Fujimoto, Scott, Lin, Hsiu-Chin, Meger, David
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911734985392128
author Houssaini, Anas
Danesh, Mohamad H.
Abyaneh, Amin
Fujimoto, Scott
Lin, Hsiu-Chin
Meger, David
author_facet Houssaini, Anas
Danesh, Mohamad H.
Abyaneh, Amin
Fujimoto, Scott
Lin, Hsiu-Chin
Meger, David
contents Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies handle this trade-off by modeling the behavior distribution to regularize the RL objective, but they require iterative denoising, solver integrations, and in more efficient variants, distillation or other approximations at inference. We propose DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement. The value signal biases the policy toward high-value regions of the data support, while attraction and repulsion together keep generated actions near the data and prevent collapse onto a single mode. DriftQL is implemented as a single network with a unified training objective and generates actions in a single forward pass. On D4RL and OGBench, DriftQL consistently outperforms diffusion and flow methods, advancing the state of the art. Under degraded data quality, where the baselines visibly struggle, DriftQL remains close to its clean-data performance, positioning it as a promising alternative to diffusion and flow-based methods while maintaining the simplicity and efficiency of deterministic approaches. Project page: https://driftql.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2606_00350
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Drift Q-Learning
Houssaini, Anas
Danesh, Mohamad H.
Abyaneh, Amin
Fujimoto, Scott
Lin, Hsiu-Chin
Meger, David
Machine Learning
Artificial Intelligence
Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies handle this trade-off by modeling the behavior distribution to regularize the RL objective, but they require iterative denoising, solver integrations, and in more efficient variants, distillation or other approximations at inference. We propose DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement. The value signal biases the policy toward high-value regions of the data support, while attraction and repulsion together keep generated actions near the data and prevent collapse onto a single mode. DriftQL is implemented as a single network with a unified training objective and generates actions in a single forward pass. On D4RL and OGBench, DriftQL consistently outperforms diffusion and flow methods, advancing the state of the art. Under degraded data quality, where the baselines visibly struggle, DriftQL remains close to its clean-data performance, positioning it as a promising alternative to diffusion and flow-based methods while maintaining the simplicity and efficiency of deterministic approaches. Project page: https://driftql.github.io/
title Drift Q-Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2606.00350