BPL: Bias-adaptive Preference Distillation Learning for Recommender System

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kang, SeongKu, Lian, Jianxun, Lee, Dongha, Kweon, Wonbin, Jang, Sanghwan, Lee, Jaehyun, Wang, Jindong, Xie, Xing, Yu, Hwanjo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917023771000832
author Kang, SeongKu
Lian, Jianxun
Lee, Dongha
Kweon, Wonbin
Jang, Sanghwan
Lee, Jaehyun
Wang, Jindong
Xie, Xing
Yu, Hwanjo
author_facet Kang, SeongKu
Lian, Jianxun
Lee, Dongha
Kweon, Wonbin
Jang, Sanghwan
Lee, Jaehyun
Wang, Jindong
Xie, Xing
Yu, Hwanjo
contents Recommender systems suffer from biases that cause the collected feedback to incompletely reveal user preference. While debiasing learning has been extensively studied, they mostly focused on the specialized (called counterfactual) test environment simulated by random exposure of items, significantly degrading accuracy in the typical (called factual) test environment based on actual user-item interactions. In fact, each test environment highlights the benefit of a different aspect: the counterfactual test emphasizes user satisfaction in the long-terms, while the factual test focuses on predicting subsequent user behaviors on platforms. Therefore, it is desirable to have a model that performs well on both tests rather than only one. In this work, we introduce a new learning framework, called Bias-adaptive Preference distillation Learning (BPL), to gradually uncover user preferences with dual distillation strategies. These distillation strategies are designed to drive high performance in both factual and counterfactual test environments. Employing a specialized form of teacher-student distillation from a biased model, BPL retains accurate preference knowledge aligned with the collected feedback, leading to high performance in the factual test. Furthermore, through self-distillation with reliability filtering, BPL iteratively refines its knowledge throughout the training process. This enables the model to produce more accurate predictions across a broader range of user-item combinations, thereby improving performance in the counterfactual test. Comprehensive experiments validate the effectiveness of BPL in both factual and counterfactual tests. Our implementation is accessible via: https://github.com/SeongKu-Kang/BPL.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16076
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BPL: Bias-adaptive Preference Distillation Learning for Recommender System
Kang, SeongKu
Lian, Jianxun
Lee, Dongha
Kweon, Wonbin
Jang, Sanghwan
Lee, Jaehyun
Wang, Jindong
Xie, Xing
Yu, Hwanjo
Machine Learning
Artificial Intelligence
Information Retrieval
Recommender systems suffer from biases that cause the collected feedback to incompletely reveal user preference. While debiasing learning has been extensively studied, they mostly focused on the specialized (called counterfactual) test environment simulated by random exposure of items, significantly degrading accuracy in the typical (called factual) test environment based on actual user-item interactions. In fact, each test environment highlights the benefit of a different aspect: the counterfactual test emphasizes user satisfaction in the long-terms, while the factual test focuses on predicting subsequent user behaviors on platforms. Therefore, it is desirable to have a model that performs well on both tests rather than only one. In this work, we introduce a new learning framework, called Bias-adaptive Preference distillation Learning (BPL), to gradually uncover user preferences with dual distillation strategies. These distillation strategies are designed to drive high performance in both factual and counterfactual test environments. Employing a specialized form of teacher-student distillation from a biased model, BPL retains accurate preference knowledge aligned with the collected feedback, leading to high performance in the factual test. Furthermore, through self-distillation with reliability filtering, BPL iteratively refines its knowledge throughout the training process. This enables the model to produce more accurate predictions across a broader range of user-item combinations, thereby improving performance in the counterfactual test. Comprehensive experiments validate the effectiveness of BPL in both factual and counterfactual tests. Our implementation is accessible via: https://github.com/SeongKu-Kang/BPL.
title BPL: Bias-adaptive Preference Distillation Learning for Recommender System
topic Machine Learning
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2510.16076