How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Cai, Hongyi James, Wang, Junlin, Chen, Xiaoyin, Dhingra, Bhuwan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
por: Huang, Yukun, et al.
Publicado: (2024)
por: Huang, Yukun, et al.
Publicado: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
por: Chen, Sanxing, et al.
Publicado: (2025)
por: Chen, Sanxing, et al.
Publicado: (2025)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
Adversarial Math Word Problem Generation
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
Staircase Streaming for Low-Latency Multi-Agent Inference
por: Wang, Junlin, et al.
Publicado: (2025)
por: Wang, Junlin, et al.
Publicado: (2025)
Real-time Factuality Assessment from Adversarial Feedback
por: Chen, Sanxing, et al.
Publicado: (2024)
por: Chen, Sanxing, et al.
Publicado: (2024)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
por: Holsman, Maximilian, et al.
Publicado: (2025)
por: Holsman, Maximilian, et al.
Publicado: (2025)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
por: Limozin, Alexis, et al.
Publicado: (2026)
por: Limozin, Alexis, et al.
Publicado: (2026)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
por: Matsutani, Kohsei, et al.
Publicado: (2025)
por: Matsutani, Kohsei, et al.
Publicado: (2025)
How Much Can RAG Help the Reasoning of LLM?
por: Liu, Jingyu, et al.
Publicado: (2024)
por: Liu, Jingyu, et al.
Publicado: (2024)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
por: Qiu, Haibo, et al.
Publicado: (2025)
por: Qiu, Haibo, et al.
Publicado: (2025)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
Coding Agents are Effective Long-Context Processors
por: Cao, Weili, et al.
Publicado: (2026)
por: Cao, Weili, et al.
Publicado: (2026)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
por: Wang, Junlin, et al.
Publicado: (2025)
por: Wang, Junlin, et al.
Publicado: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
por: Li, Youpeng, et al.
Publicado: (2026)
por: Li, Youpeng, et al.
Publicado: (2026)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
por: Liu, Zihan, et al.
Publicado: (2025)
por: Liu, Zihan, et al.
Publicado: (2025)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
por: Wu, Qinzhuo, et al.
Publicado: (2025)
por: Wu, Qinzhuo, et al.
Publicado: (2025)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
por: Zhao, Yang, et al.
Publicado: (2026)
por: Zhao, Yang, et al.
Publicado: (2026)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
por: Huang, Yukun, et al.
Publicado: (2025)
por: Huang, Yukun, et al.
Publicado: (2025)
RL Fine-Tuning Heals OOD Forgetting in SFT
por: Jin, Hangzhan, et al.
Publicado: (2025)
por: Jin, Hangzhan, et al.
Publicado: (2025)
Document-as-Image Representations Fall Short for Scientific Retrieval
por: Khalighinejad, Ghazal, et al.
Publicado: (2026)
por: Khalighinejad, Ghazal, et al.
Publicado: (2026)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
por: Wei, Lai, et al.
Publicado: (2025)
por: Wei, Lai, et al.
Publicado: (2025)
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
por: Rezazadeh, Navid, et al.
Publicado: (2026)
por: Rezazadeh, Navid, et al.
Publicado: (2026)
Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
por: Zhang, Minxing, et al.
Publicado: (2025)
por: Zhang, Minxing, et al.
Publicado: (2025)
Emergent Search and Backtracking in Latent Reasoning Models
por: Cui, Jasmine, et al.
Publicado: (2026)
por: Cui, Jasmine, et al.
Publicado: (2026)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
por: Wang, Sudong, et al.
Publicado: (2026)
por: Wang, Sudong, et al.
Publicado: (2026)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
por: Kang, Feiyang, et al.
Publicado: (2025)
por: Kang, Feiyang, et al.
Publicado: (2025)
Calibrating Long-form Generations from Large Language Models
por: Huang, Yukun, et al.
Publicado: (2024)
por: Huang, Yukun, et al.
Publicado: (2024)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
por: Ge, Yuyao, et al.
Publicado: (2025)
por: Ge, Yuyao, et al.
Publicado: (2025)
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
por: Kim, Joongwon, et al.
Publicado: (2025)
por: Kim, Joongwon, et al.
Publicado: (2025)
How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure
por: Thompson, Paul M.
Publicado: (2025)
por: Thompson, Paul M.
Publicado: (2025)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
por: Wang, Xiaoyin, et al.
Publicado: (2024)
por: Wang, Xiaoyin, et al.
Publicado: (2024)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
por: Ren, Qihan, et al.
Publicado: (2026)
por: Ren, Qihan, et al.
Publicado: (2026)
Backtracking for Safety
por: Sel, Bilgehan, et al.
Publicado: (2025)
por: Sel, Bilgehan, et al.
Publicado: (2025)
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
por: Huang, Yukun, et al.
Publicado: (2026)
por: Huang, Yukun, et al.
Publicado: (2026)
How Much Data is Enough? The Zeta Law of Discoverability in Biomedical Data, featuring the enigmatic Riemann zeta function
por: Thompson, Paul M.
Publicado: (2026)
por: Thompson, Paul M.
Publicado: (2026)
Ejemplares similares
-
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
por: Huang, Yukun, et al.
Publicado: (2024) -
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
por: Chen, Sanxing, et al.
Publicado: (2025) -
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
por: Zhai, Zhiyuan, et al.
Publicado: (2026) -
Adversarial Math Word Problem Generation
por: Xie, Roy, et al.
Publicado: (2024) -
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)