RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Matsutani, Kohsei, Takashiro, Shota, Minegishi, Gouki, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026)
by: Matsutani, Kohsei, et al.
Published: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026)
by: Feng, Jingyuan, et al.
Published: (2026)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
by: Watanabe, Yusuke, et al.
Published: (2026)
by: Watanabe, Yusuke, et al.
Published: (2026)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
by: Takashiro, Shota, et al.
Published: (2024)
by: Takashiro, Shota, et al.
Published: (2024)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
by: Uchiyama, Fumiya, et al.
Published: (2024)
by: Uchiyama, Fumiya, et al.
Published: (2024)
Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
by: Takagi, Hirohane, et al.
Published: (2025)
by: Takagi, Hirohane, et al.
Published: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs
by: Chen, Yingjian, et al.
Published: (2026)
by: Chen, Yingjian, et al.
Published: (2026)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
by: Kim, Bum Jun, et al.
Published: (2026)
by: Kim, Bum Jun, et al.
Published: (2026)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
by: Kojima, Takeshi, et al.
Published: (2025)
by: Kojima, Takeshi, et al.
Published: (2025)
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4
by: Guo, Jiaxian, et al.
Published: (2023)
by: Guo, Jiaxian, et al.
Published: (2023)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
by: Kubo, Kenji, et al.
Published: (2026)
by: Kubo, Kenji, et al.
Published: (2026)
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
by: Sendai, Kohei, et al.
Published: (2025)
by: Sendai, Kohei, et al.
Published: (2025)
Steering at the Source: Style Modulation Heads for Robust Persona Control
by: Izawa, Yoshihiro, et al.
Published: (2026)
by: Izawa, Yoshihiro, et al.
Published: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
by: Zhao, Xinjie, et al.
Published: (2025)
by: Zhao, Xinjie, et al.
Published: (2025)
Dynamic Injection of Entity Knowledge into Dense Retrievers
by: Yamada, Ikuya, et al.
Published: (2025)
by: Yamada, Ikuya, et al.
Published: (2025)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
by: Harada, Keno, et al.
Published: (2025)
by: Harada, Keno, et al.
Published: (2025)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
by: Kojima, Takeshi, et al.
Published: (2024)
by: Kojima, Takeshi, et al.
Published: (2024)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
by: Cai, Hongyi James, et al.
Published: (2025)
by: Cai, Hongyi James, et al.
Published: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
CLIP-like Model as a Foundational Density Ratio Estimator
by: Uchiyama, Fumiya, et al.
Published: (2025)
by: Uchiyama, Fumiya, et al.
Published: (2025)
RL Fine-Tuning Heals OOD Forgetting in SFT
by: Jin, Hangzhan, et al.
Published: (2025)
by: Jin, Hangzhan, et al.
Published: (2025)
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
by: Kawakami, Wataru, et al.
Published: (2025)
by: Kawakami, Wataru, et al.
Published: (2025)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Similar Items
-
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026) -
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
by: Minegishi, Gouki, et al.
Published: (2025) -
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026) -
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026) -
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)