Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Mingqiao, Peng, Qiyao, Wang, Yinghui, Liu, Hongtao, Wang, Yumeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916048598466560
author Zhang, Mingqiao
Peng, Qiyao
Wang, Yinghui
Liu, Hongtao
Wang, Yumeng
author_facet Zhang, Mingqiao
Peng, Qiyao
Wang, Yinghui
Liu, Hongtao
Wang, Yumeng
contents The expanding integration of Large Language Models (LLMs) into recommender systems poses critical challenges to evaluation reliability. This paper identifies and investigates a previously overlooked issue: benchmark data leakage in LLM-based recommendation. This phenomenon occurs when LLMs are exposed to and potentially memorize benchmark datasets during pre-training or fine-tuning, leading to artificially inflated performance metrics that fail to reflect true model performance. To validate this phenomenon, we simulate diverse data leakage scenarios by conducting continued pre-training of foundation models on strategically blended corpora, which include user-item interactions from both in-domain and out-of-domain sources. Our experiments reveal a dual-effect of data leakage: when the leaked data is domain-relevant, it induces substantial but spurious performance gains, misleadingly exaggerating the model's capability. In contrast, domain-irrelevant leakage typically degrades recommendation accuracy, highlighting the complex and contingent nature of this contamination. Our findings reveal that data leakage acts as a critical, previously unaccounted-for factor in LLM-based recommendation, which could impact the true model performance. We release our code at https://github.com/yusba1/LLMRec-Data-Leakage.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13626
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?
Zhang, Mingqiao
Peng, Qiyao
Wang, Yinghui
Liu, Hongtao
Wang, Yumeng
Machine Learning
The expanding integration of Large Language Models (LLMs) into recommender systems poses critical challenges to evaluation reliability. This paper identifies and investigates a previously overlooked issue: benchmark data leakage in LLM-based recommendation. This phenomenon occurs when LLMs are exposed to and potentially memorize benchmark datasets during pre-training or fine-tuning, leading to artificially inflated performance metrics that fail to reflect true model performance. To validate this phenomenon, we simulate diverse data leakage scenarios by conducting continued pre-training of foundation models on strategically blended corpora, which include user-item interactions from both in-domain and out-of-domain sources. Our experiments reveal a dual-effect of data leakage: when the leaked data is domain-relevant, it induces substantial but spurious performance gains, misleadingly exaggerating the model's capability. In contrast, domain-irrelevant leakage typically degrades recommendation accuracy, highlighting the complex and contingent nature of this contamination. Our findings reveal that data leakage acts as a critical, previously unaccounted-for factor in LLM-based recommendation, which could impact the true model performance. We release our code at https://github.com/yusba1/LLMRec-Data-Leakage.
title Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?
topic Machine Learning
url https://arxiv.org/abs/2602.13626