On the Reliability of Sampling Strategies in Offline Recommender Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pereira, Bruno L., Said, Alan, Santos, Rodrygo L. T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915439035023360
author Pereira, Bruno L.
Said, Alan
Santos, Rodrygo L. T.
author_facet Pereira, Bruno L.
Said, Alan
Santos, Rodrygo L. T.
contents Offline evaluation plays a central role in benchmarking recommender systems when online testing is impractical or risky. However, it is susceptible to two key sources of bias: exposure bias, where users only interact with items they are shown, and sampling bias, introduced when evaluation is performed on a subset of logged items rather than the full catalog. While prior work has proposed methods to mitigate sampling bias, these are typically assessed on fixed logged datasets rather than for their ability to support reliable model comparisons under varying exposure conditions or relative to true user preferences. In this paper, we investigate how different combinations of logging and sampling choices affect the reliability of offline evaluation. Using a fully observed dataset as ground truth, we systematically simulate diverse exposure biases and assess the reliability of common sampling strategies along four dimensions: sampling resolution (recommender model separability), fidelity (agreement with full evaluation), robustness (stability under exposure bias), and predictive power (alignment with ground truth). Our findings highlight when and how sampling distorts evaluation outcomes and offer practical guidance for selecting strategies that yield faithful and robust offline comparisons.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Reliability of Sampling Strategies in Offline Recommender Evaluation
Pereira, Bruno L.
Said, Alan
Santos, Rodrygo L. T.
Information Retrieval
Machine Learning
Offline evaluation plays a central role in benchmarking recommender systems when online testing is impractical or risky. However, it is susceptible to two key sources of bias: exposure bias, where users only interact with items they are shown, and sampling bias, introduced when evaluation is performed on a subset of logged items rather than the full catalog. While prior work has proposed methods to mitigate sampling bias, these are typically assessed on fixed logged datasets rather than for their ability to support reliable model comparisons under varying exposure conditions or relative to true user preferences. In this paper, we investigate how different combinations of logging and sampling choices affect the reliability of offline evaluation. Using a fully observed dataset as ground truth, we systematically simulate diverse exposure biases and assess the reliability of common sampling strategies along four dimensions: sampling resolution (recommender model separability), fidelity (agreement with full evaluation), robustness (stability under exposure bias), and predictive power (alignment with ground truth). Our findings highlight when and how sampling distorts evaluation outcomes and offer practical guidance for selecting strategies that yield faithful and robust offline comparisons.
title On the Reliability of Sampling Strategies in Offline Recommender Evaluation
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2508.05398