How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Hongyi James, Wang, Junlin, Chen, Xiaoyin, Dhingra, Bhuwan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
by: Chen, Sanxing, et al.
Published: (2025)
by: Chen, Sanxing, et al.
Published: (2025)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
Staircase Streaming for Low-Latency Multi-Agent Inference
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Real-time Factuality Assessment from Adversarial Feedback
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
by: Holsman, Maximilian, et al.
Published: (2025)
by: Holsman, Maximilian, et al.
Published: (2025)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
by: Matsutani, Kohsei, et al.
Published: (2025)
by: Matsutani, Kohsei, et al.
Published: (2025)
How Much Can RAG Help the Reasoning of LLM?
by: Liu, Jingyu, et al.
Published: (2024)
by: Liu, Jingyu, et al.
Published: (2024)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
Coding Agents are Effective Long-Context Processors
by: Cao, Weili, et al.
Published: (2026)
by: Cao, Weili, et al.
Published: (2026)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2026)
by: Li, Youpeng, et al.
Published: (2026)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
by: Wu, Qinzhuo, et al.
Published: (2025)
by: Wu, Qinzhuo, et al.
Published: (2025)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
RL Fine-Tuning Heals OOD Forgetting in SFT
by: Jin, Hangzhan, et al.
Published: (2025)
by: Jin, Hangzhan, et al.
Published: (2025)
Document-as-Image Representations Fall Short for Scientific Retrieval
by: Khalighinejad, Ghazal, et al.
Published: (2026)
by: Khalighinejad, Ghazal, et al.
Published: (2026)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
by: Zhang, Minxing, et al.
Published: (2025)
by: Zhang, Minxing, et al.
Published: (2025)
Emergent Search and Backtracking in Latent Reasoning Models
by: Cui, Jasmine, et al.
Published: (2026)
by: Cui, Jasmine, et al.
Published: (2026)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Calibrating Long-form Generations from Large Language Models
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
by: Kim, Joongwon, et al.
Published: (2025)
by: Kim, Joongwon, et al.
Published: (2025)
How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure
by: Thompson, Paul M.
Published: (2025)
by: Thompson, Paul M.
Published: (2025)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
by: Wang, Xiaoyin, et al.
Published: (2024)
by: Wang, Xiaoyin, et al.
Published: (2024)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
Backtracking for Safety
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
by: Huang, Yukun, et al.
Published: (2026)
by: Huang, Yukun, et al.
Published: (2026)
How Much Data is Enough? The Zeta Law of Discoverability in Biomedical Data, featuring the enigmatic Riemann zeta function
by: Thompson, Paul M.
Published: (2026)
by: Thompson, Paul M.
Published: (2026)
Similar Items
-
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
by: Huang, Yukun, et al.
Published: (2024) -
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
by: Chen, Sanxing, et al.
Published: (2025) -
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
by: Zhai, Zhiyuan, et al.
Published: (2026) -
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024) -
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)