Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Jiang, Nan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910200122834944
author Li, Xiang
Jiang, Nan
author_facet Li, Xiang
Jiang, Nan
contents We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximate the expected return under the target policy. The weights are learned inductively in a top-down manner via a moment matching objective against a value-function discriminator class. Notably, and perhaps surprisingly, a data-dependent finite-sample guarantee for general function approximation can be established under only the realizability of $Q^π$, with a dimension-free bound -- that is, the error does not depend on the statistical complexity of the function class. We also establish connections to several existing methods, such as importance sampling and linear FQE. Further theoretical analyses shed new light on the nature of coverage, a concept of fundamental importance to offline RL.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06474
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
Li, Xiang
Jiang, Nan
Machine Learning
Artificial Intelligence
We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximate the expected return under the target policy. The weights are learned inductively in a top-down manner via a moment matching objective against a value-function discriminator class. Notably, and perhaps surprisingly, a data-dependent finite-sample guarantee for general function approximation can be established under only the realizability of $Q^π$, with a dimension-free bound -- that is, the error does not depend on the statistical complexity of the function class. We also establish connections to several existing methods, such as importance sampling and linear FQE. Further theoretical analyses shed new light on the nature of coverage, a concept of fundamental importance to offline RL.
title Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.06474