arxiv:2507.10532

Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination

Published on Jul 14

· Submitted by

zsytony on Jul 15

#1 Paper of the day

Upvote

Authors:

Songyang Zhang ,

Abstract

Research on enhancing LLM reasoning through RL reveals that accurate reward signals are crucial for performance improvement, and current benchmarks may be unreliable due to data contamination.

AI-generated summary

The reasoning capabilities of large language models (LLMs) have been a longstanding focus of research. Recent works have further enhanced these capabilities using reinforcement learning (RL), with many new methods claiming significant improvements with minimal or no external supervision. Surprisingly, some studies even suggest that random or incorrect reward signals can enhance reasoning performance. However, these breakthroughs are mostly reported on the Qwen2.5 model family and evaluated on well-known benchmarks such as MATH-500, AMC, and AIME, while failing to achieve similar gains on other models like Llama, which warrants further investigation. Our analysis shows that although Qwen2.5 achieves strong mathematical reasoning performance, its pretraining on large-scale web corpora makes it vulnerable to data contamination in popular benchmarks. As a result, results derived from these benchmarks may be unreliable. To address this, we introduce a generator that produces fully synthetic arithmetic problems of arbitrary length and difficulty, yielding a clean dataset we call RandomCalculation. Using these leakage-free datasets, we show that only accurate reward signals consistently improve performance, while noisy or incorrect signals do not. We advocate for evaluating RL methods on uncontaminated benchmarks and across diverse model families to ensure trustworthy conclusions.

View arXiv page View PDF Add to collection

Community

zsytony

Paper author Paper submitter 2 days ago

Report

guanning

about 12 hours ago

•

edited about 12 hours ago

Thanks very much for this great work!
Are you planning to upload the RandomCalculation benchmark in the near recent? (So that we can evaluate on more models lol)🙂