Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Mingqi, Zhang, Zhihao, Dong, Qiaole, Xi, Zhiheng, Zhao, Jun, Jin, Senjie, Fan, Xiaoran, Zhou, Yuhao, Lv, Huijie, Zhang, Ming, Fu, Yanwei, Liu, Qin, Zhang, Songyang, Zhang, Qi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914204930277376
author Wu, Mingqi
Zhang, Zhihao
Dong, Qiaole
Xi, Zhiheng
Zhao, Jun
Jin, Senjie
Fan, Xiaoran
Zhou, Yuhao
Lv, Huijie
Zhang, Ming
Fu, Yanwei
Liu, Qin
Zhang, Songyang
Zhang, Qi
author_facet Wu, Mingqi
Zhang, Zhihao
Dong, Qiaole
Xi, Zhiheng
Zhao, Jun
Jin, Senjie
Fan, Xiaoran
Zhou, Yuhao
Lv, Huijie
Zhang, Ming
Fu, Yanwei
Liu, Qin
Zhang, Songyang
Zhang, Qi
contents Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model's performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
Wu, Mingqi
Zhang, Zhihao
Dong, Qiaole
Xi, Zhiheng
Zhao, Jun
Jin, Senjie
Fan, Xiaoran
Zhou, Yuhao
Lv, Huijie
Zhang, Ming
Fu, Yanwei
Liu, Qin
Zhang, Songyang
Zhang, Qi
Machine Learning
Artificial Intelligence
Computation and Language
Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield substantial performance gains with minimal or even no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance performance. However, these breakthroughs are predominantly observed for the mathematically strong Qwen2.5 series on benchmarks such as MATH-500, AMC, and AIME, and seldom transfer to models like Llama, which warrants a more in-depth investigation. In this work, our empirical analysis reveals that pre-training on massive web-scale corpora leaves Qwen2.5 susceptible to data contamination in widely used benchmarks. Consequently, conclusions derived from contaminated benchmarks on Qwen2.5 series may be unreliable. To obtain trustworthy evaluation results, we introduce a generator that creates fully clean arithmetic problems of arbitrary length and difficulty, dubbed RandomCalculation. Using this leakage-free dataset, we show that only accurate reward signals yield steady improvements that surpass the base model's performance boundary in mathematical reasoning, whereas random or incorrect rewards do not. Moreover, we conduct more fine-grained analyses to elucidate the factors underlying the different performance observed on the MATH-500 and RandomCalculation benchmarks. Consequently, we recommend that future studies evaluate models on uncontaminated benchmarks and, when feasible, test various model series to ensure trustworthy conclusions about RL and related methods.
title Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.10532