Saved in:
Bibliographic Details
Main Authors: Zhang, Yichi, Belloni, Alexander, Fang, Ethan X., Lu, Junwei, Xu, Xiaoan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.05852
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916937784623104
author Zhang, Yichi
Belloni, Alexander
Fang, Ethan X.
Lu, Junwei
Xu, Xiaoan
author_facet Zhang, Yichi
Belloni, Alexander
Fang, Ethan X.
Lu, Junwei
Xu, Xiaoan
contents Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-dependent preference score functions across domains. Focusing on the contextual Bradley-Terry-Luce model, we develop a semiparametric efficient estimator that automates the debiased estimation through aggregating weighted residual balancing terms across the comparison graph. We show that the efficiency is achieved when the weights are derived from a novel strategy called Fisher random walk. We also propose a computationally feasible method to compute the weights by a potential representation of nuisance weight functions. We show our inference procedure is valid for general score function estimators accommodating the practitioners' need to implement flexible deep learning methods. We extend the procedure to multiple hypothesis testing using a Gaussian multiplier bootstrap that controls familywise error and to distributional shift via a cross-fitted importance-sampling adjustment for target-domain inference. Numerical studies, including language model evaluations under diverse contexts, corroborate the accuracy, efficiency, and practical utility of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
Zhang, Yichi
Belloni, Alexander
Fang, Ethan X.
Lu, Junwei
Xu, Xiaoan
Machine Learning
Statistics Theory
Motivated by the need for rigorous and scalable evaluation of large language models, we study contextual preference inference for pairwise comparison functionals of context-dependent preference score functions across domains. Focusing on the contextual Bradley-Terry-Luce model, we develop a semiparametric efficient estimator that automates the debiased estimation through aggregating weighted residual balancing terms across the comparison graph. We show that the efficiency is achieved when the weights are derived from a novel strategy called Fisher random walk. We also propose a computationally feasible method to compute the weights by a potential representation of nuisance weight functions. We show our inference procedure is valid for general score function estimators accommodating the practitioners' need to implement flexible deep learning methods. We extend the procedure to multiple hypothesis testing using a Gaussian multiplier bootstrap that controls familywise error and to distributional shift via a cross-fitted importance-sampling adjustment for target-domain inference. Numerical studies, including language model evaluations under diverse contexts, corroborate the accuracy, efficiency, and practical utility of our method.
title Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2509.05852