Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Minwu, Shrestha, Anubhav, Shrestha, Safal, Nepal, Aadim, Ross, Keith
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914125778518016
author Kim, Minwu
Shrestha, Anubhav
Shrestha, Safal
Nepal, Aadim
Ross, Keith
author_facet Kim, Minwu
Shrestha, Anubhav
Shrestha, Safal
Nepal, Aadim
Ross, Keith
contents Recent studies have shown that reinforcement learning with verifiable rewards (RLVR) enhances overall accuracy (pass@1) but often fails to improve capability (pass@k) of LLMs in reasoning tasks, while distillation can improve both. In this paper, we investigate the mechanisms behind these phenomena. First, we demonstrate that RLVR struggles to improve capability as it focuses on improving the accuracy of the easier questions to the detriment of the accuracy of the most difficult questions. Second, we show that RLVR does not merely increase the success probability for the easier questions, but in our small model settings, produces quality responses that were absent in its original output distribution. In addition, we show these responses are neither noticeably longer nor feature more reflection-related keywords, underscoring the need for more reliable indicators of response quality. Third, from the experiment distilling teacher responses to in-distribution problems, we find that capability does not always improve with distillation. We conjecture that capability improves only when new knowledge is introduced, whereas distilling reasoning patterns only improves accuracy but not capability, sacrificing performance on the most difficult questions, similar to RLVR. Together, these findings offer a clearer understanding of how RLVR and distillation shape reasoning behavior in LLMs
format Preprint
id arxiv_https___arxiv_org_abs_2505_14216
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
Kim, Minwu
Shrestha, Anubhav
Shrestha, Safal
Nepal, Aadim
Ross, Keith
Artificial Intelligence
Computation and Language
Recent studies have shown that reinforcement learning with verifiable rewards (RLVR) enhances overall accuracy (pass@1) but often fails to improve capability (pass@k) of LLMs in reasoning tasks, while distillation can improve both. In this paper, we investigate the mechanisms behind these phenomena. First, we demonstrate that RLVR struggles to improve capability as it focuses on improving the accuracy of the easier questions to the detriment of the accuracy of the most difficult questions. Second, we show that RLVR does not merely increase the success probability for the easier questions, but in our small model settings, produces quality responses that were absent in its original output distribution. In addition, we show these responses are neither noticeably longer nor feature more reflection-related keywords, underscoring the need for more reliable indicators of response quality. Third, from the experiment distilling teacher responses to in-distribution problems, we find that capability does not always improve with distillation. We conjecture that capability improves only when new knowledge is introduced, whereas distilling reasoning patterns only improves accuracy but not capability, sacrificing performance on the most difficult questions, similar to RLVR. Together, these findings offer a clearer understanding of how RLVR and distillation shape reasoning behavior in LLMs
title Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.14216