Meta Flow Maps enable scalable reward alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Potaptchik, Peter, Saravanan, Adhi, Mammadov, Abbas, Prat, Alvaro, Albergo, Michael S., Teh, Yee Whye
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915742576803840
author Potaptchik, Peter
Saravanan, Adhi
Mammadov, Abbas
Prat, Alvaro
Albergo, Michael S.
Teh, Yee Whye
author_facet Potaptchik, Peter
Saravanan, Adhi
Mammadov, Abbas
Prat, Alvaro
Albergo, Michael S.
Teh, Yee Whye
contents Controlling generative models is computationally expensive. This is because optimal alignment with a reward function--whether via inference-time steering or fine-tuning--requires estimating the value function. This task demands access to the conditional posterior $p_{1|t}(x_1|x_t)$, the distribution of clean data $x_1$ consistent with an intermediate state $x_t$, a requirement that typically compels methods to resort to costly trajectory simulations. To address this bottleneck, we introduce Meta Flow Maps (MFMs), a framework extending consistency models and flow maps into the stochastic regime. MFMs are trained to perform stochastic one-step posterior sampling, generating arbitrarily many i.i.d. draws of clean data $x_1$ from any intermediate state. Crucially, these samples provide a differentiable reparametrization that unlocks efficient value function estimation. We leverage this capability to solve bottlenecks in both paradigms: enabling inference-time steering without inner rollouts, and facilitating unbiased, off-policy fine-tuning to general rewards. Empirically, our single-particle steered-MFM sampler outperforms a Best-of-1000 baseline on ImageNet across multiple rewards at a fraction of the compute.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14430
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Meta Flow Maps enable scalable reward alignment
Potaptchik, Peter
Saravanan, Adhi
Mammadov, Abbas
Prat, Alvaro
Albergo, Michael S.
Teh, Yee Whye
Machine Learning
Controlling generative models is computationally expensive. This is because optimal alignment with a reward function--whether via inference-time steering or fine-tuning--requires estimating the value function. This task demands access to the conditional posterior $p_{1|t}(x_1|x_t)$, the distribution of clean data $x_1$ consistent with an intermediate state $x_t$, a requirement that typically compels methods to resort to costly trajectory simulations. To address this bottleneck, we introduce Meta Flow Maps (MFMs), a framework extending consistency models and flow maps into the stochastic regime. MFMs are trained to perform stochastic one-step posterior sampling, generating arbitrarily many i.i.d. draws of clean data $x_1$ from any intermediate state. Crucially, these samples provide a differentiable reparametrization that unlocks efficient value function estimation. We leverage this capability to solve bottlenecks in both paradigms: enabling inference-time steering without inner rollouts, and facilitating unbiased, off-policy fine-tuning to general rewards. Empirically, our single-particle steered-MFM sampler outperforms a Best-of-1000 baseline on ImageNet across multiple rewards at a fraction of the compute.
title Meta Flow Maps enable scalable reward alignment
topic Machine Learning
url https://arxiv.org/abs/2601.14430