"Does RL really incentivize reasoning beyond the base model?"

At large numbers of samples, base models solve problems that their RL-trained versions do not, so RLVR appears to sharpen what the base model can already do and to add few new reasoning abilities.

Date
18 April 2025
Who
Tsinghua University LeapLab, Shanghai Jiao Tong University; counter-evidence from NVIDIA and the Allen Institute for AI
People
Yang Yue, Zhiqi Chen, Gao Huang
Confidence
High on the papers; the interpretation is contested
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): Tsinghua University LeapLab, Shanghai Jiao Tong University; counter-evidence from NVIDIA and the Allen Institute for AI · People: Yang Yue, Zhiqi Chen, Gao Huang · Confidence: High on the papers; the interpretation is contested Primary sources: arXiv 2504.13837 (v1 2025-04-18; v5 2025-11-24) · ProRL, arXiv 2505.24864 · Spurious Rewards, arXiv 2506.10947

One-liner. At large numbers of samples, base models solve problems that their RL-trained versions do not, so RLVR appears to sharpen what the base model can already do and to add few new reasoning abilities.

Why it happened. After R1 the field assumed RL with verifiable rewards (RLVR) would self-improve like AlphaZero. A Tsinghua group tested the claim directly with pass@k at large k, which measures whether any of k attempts is right, alongside pass@1.

The idea. If RL created new reasoning ability, an RL-trained model's pass@k curve should stay above the base model's at all k. Across model families, six RLVR algorithms, and math, coding and visual benchmarks, the authors found RLVR models win at small k but base models catch up and exceed them at large k; the reasoning boundary "often narrows as RLVR training progresses"; the RL model's reasoning paths already lie within the base model's sampling distribution (coverage and perplexity analysis). By contrast, distillation from a stronger teacher can introduce new patterns. The six RL algorithms performed similarly and fell far short of using the base model's full potential (paper).

Results and counter-evidence. NVIDIA's ProRL (arXiv 2025-05-30) argued the finding reflects short RL training. With KL control, reference-policy resets and diverse tasks over prolonged training, a 1.5B model outperformed its base at all pass@k, including on tasks the base failed at (ProRL). The "Spurious Rewards" paper (2025-06-12; revised 2026-02-25) found random or incorrect reward signals raised Qwen2.5-Math-7B's MATH-500 score by 21.4 points against 29.1 for correct rewards, attributed to a GRPO clipping bias that amplifies pre-trained behaviour (such as code-style reasoning), and that this did not carry over to Llama or OLMo (paper). Results obtained only on Qwen bases may not generalise.

How it spread. I infer that the debate coincided with a 2025 to 2026 shift toward scaling RL compute (B08-40) and richer rewards such as proof verifiers (DeepSeekMath-V2), and that its finding that distillation adds capability is consistent with labs' concern about distillation (B08-14); the paper is not cited as the cause.

Why it mattered. It gave the field a precise way to say what RLVR does, which is to raise reliability on the problems the base model can already sometimes solve. Whether it can discover new strategies at larger scale is the central open question since R1 (B24).

Nuance, controversy and myths. "RL doesn't teach anything new" overstates a finding about pass@k on specific benchmarks and models at modest RL scale; "RL discovers new reasoning" overstates the "aha moment." Sharpening is demonstrated, and expansion is plausible at scale but not settled. I infer that large-k pass@k can also reward lucky guesses on tasks with short or constrained answers, so the metric is best read with proof-style or code tasks.

Interview kit.

  • 30-second version: A Tsinghua team showed that with enough samples the base model beats its RL-tuned version, so RLVR mostly makes the model reliably find solutions it could already occasionally reach; follow-ups show longer, more diverse RL can push beyond.
  • Likely follow-ups: So is the R1 "aha" fake? → No; but pass@1 gains are not proof of new capability. What does expand capability? → Distillation from stronger teachers, bigger bases, richer RL environments.
  • Common mistake: Citing the paper as proof RL is useless; it is a sampling-efficiency claim.
  • Connect it to: B08-10, B08-40, B24.

Sources. All three papers opened (PDF of the first read locally).

Read it in the deep dive