Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minwu, Shrestha, Anubhav, Shrestha, Safal, Nepal, Aadim, Ross, Keith |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026)
by: Shrestha, Safal, et al.
Published: (2026)
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026)
by: Kim, Minwu, et al.
Published: (2026)
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)
by: Nepal, Aadim, et al.
Published: (2025)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
by: Shrestha, Manil, et al.
Published: (2025)
by: Shrestha, Manil, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
Conformal Prediction for Risk-Controlled Medical Entity Extraction Across Clinical Domains
by: Shrestha, Manil, et al.
Published: (2026)
by: Shrestha, Manil, et al.
Published: (2026)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
A Survey on LLM-Assisted Clinical Trial Recruitment
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Enabling LLM Knowledge Analysis via Extensive Materialization
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)
by: Xu, Jillian, et al.
Published: (2025)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning
by: Hu, Bokai, et al.
Published: (2024)
by: Hu, Bokai, et al.
Published: (2024)
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
by: Wu, Zhaofeng, et al.
Published: (2023)
by: Wu, Zhaofeng, et al.
Published: (2023)
Difficulty Estimation and Simplification of French Text Using LLMs
by: Jamet, Henri, et al.
Published: (2024)
by: Jamet, Henri, et al.
Published: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
by: Huang, Jiazhen, et al.
Published: (2026)
by: Huang, Jiazhen, et al.
Published: (2026)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
by: Gautam, Aayush, et al.
Published: (2025)
by: Gautam, Aayush, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
by: Luo, Zihan, et al.
Published: (2024)
by: Luo, Zihan, et al.
Published: (2024)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
by: Xie, Tian, et al.
Published: (2025)
by: Xie, Tian, et al.
Published: (2025)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
by: Zhou, Weilin, et al.
Published: (2026)
by: Zhou, Weilin, et al.
Published: (2026)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
by: Jiao, Rui, et al.
Published: (2025)
by: Jiao, Rui, et al.
Published: (2025)
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
by: Weyssow, Martin, et al.
Published: (2025)
by: Weyssow, Martin, et al.
Published: (2025)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Similar Items
-
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
by: Shrestha, Safal, et al.
Published: (2025) -
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026) -
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026) -
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025) -
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)