floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Agrawalla, Bhavya, Nauman, Michal, Agrawal, Khush, Kumar, Aviral
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912666232029184
author Agrawalla, Bhavya
Nauman, Michal
Agrawal, Khush
Kumar, Aviral
author_facet Agrawalla, Bhavya
Nauman, Michal
Agrawal, Khush
Kumar, Aviral
contents A hallmark of modern large-scale machine learning techniques is the use of training objectives that provide dense supervision to intermediate computations, such as teacher forcing the next token in language models or denoising step-by-step in diffusion models. This enables models to learn complex functions in a generalizable manner. Motivated by this observation, we investigate the benefits of iterative computation for temporal difference (TD) methods in reinforcement learning (RL). Typically they represent value functions in a monolithic fashion, without iterative compute. We introduce floq (flow-matching Q-functions), an approach that parameterizes the Q-function using a velocity field and trains it using techniques from flow-matching, typically used in generative modeling. This velocity field underneath the flow is trained using a TD-learning objective, which bootstraps from values produced by a target velocity field, computed by running multiple steps of numerical integration. Crucially, floq allows for more fine-grained control and scaling of the Q-function capacity than monolithic architectures, by appropriately setting the number of integration steps. Across a suite of challenging offline RL benchmarks and online fine-tuning tasks, floq improves performance by nearly 1.8x. floq scales capacity far better than standard TD-learning architectures, highlighting the potential of iterative computation for value learning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
Agrawalla, Bhavya
Nauman, Michal
Agrawal, Khush
Kumar, Aviral
Machine Learning
Artificial Intelligence
A hallmark of modern large-scale machine learning techniques is the use of training objectives that provide dense supervision to intermediate computations, such as teacher forcing the next token in language models or denoising step-by-step in diffusion models. This enables models to learn complex functions in a generalizable manner. Motivated by this observation, we investigate the benefits of iterative computation for temporal difference (TD) methods in reinforcement learning (RL). Typically they represent value functions in a monolithic fashion, without iterative compute. We introduce floq (flow-matching Q-functions), an approach that parameterizes the Q-function using a velocity field and trains it using techniques from flow-matching, typically used in generative modeling. This velocity field underneath the flow is trained using a TD-learning objective, which bootstraps from values produced by a target velocity field, computed by running multiple steps of numerical integration. Crucially, floq allows for more fine-grained control and scaling of the Q-function capacity than monolithic architectures, by appropriately setting the number of integration steps. Across a suite of challenging offline RL benchmarks and online fine-tuning tasks, floq improves performance by nearly 1.8x. floq scales capacity far better than standard TD-learning architectures, highlighting the potential of iterative computation for value learning.
title floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.06863