Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Gan, Xingwei, Zhu, Ying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Where does output diversity collapse in post-training?
por: Karouzos, Constantinos, et al.
Publicado: (2026)
por: Karouzos, Constantinos, et al.
Publicado: (2026)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
por: Koh, Woosung, et al.
Publicado: (2026)
por: Koh, Woosung, et al.
Publicado: (2026)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
por: Peng, Ruiying, et al.
Publicado: (2026)
por: Peng, Ruiying, et al.
Publicado: (2026)
Debunk the Myth of SFT Generalization
por: Lin, Xiaofeng, et al.
Publicado: (2025)
por: Lin, Xiaofeng, et al.
Publicado: (2025)
RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
por: Yoshihara, Hiroshi, et al.
Publicado: (2025)
por: Yoshihara, Hiroshi, et al.
Publicado: (2025)
It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt
por: Bladon, Stuart, et al.
Publicado: (2026)
por: Bladon, Stuart, et al.
Publicado: (2026)
Leveraging LLMs for reward function design in reinforcement learning control tasks
por: Cardenoso, Franklin, et al.
Publicado: (2025)
por: Cardenoso, Franklin, et al.
Publicado: (2025)
Normalization and effective learning rates in reinforcement learning
por: Lyle, Clare, et al.
Publicado: (2024)
por: Lyle, Clare, et al.
Publicado: (2024)
RL Fine-Tuning Heals OOD Forgetting in SFT
por: Jin, Hangzhan, et al.
Publicado: (2025)
por: Jin, Hangzhan, et al.
Publicado: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
Curriculum reinforcement learning with measurable task representation learning
por: Wen, Yongyan, et al.
Publicado: (2026)
por: Wen, Yongyan, et al.
Publicado: (2026)
Solving the flexible job-shop scheduling problem through an enhanced deep reinforcement learning approach
por: Echeverria, Imanol, et al.
Publicado: (2023)
por: Echeverria, Imanol, et al.
Publicado: (2023)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
por: Kang, Feiyang, et al.
Publicado: (2025)
por: Kang, Feiyang, et al.
Publicado: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
por: Chu, Tianzhe, et al.
Publicado: (2025)
por: Chu, Tianzhe, et al.
Publicado: (2025)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
por: Sukhija, Bhavya, et al.
Publicado: (2024)
por: Sukhija, Bhavya, et al.
Publicado: (2024)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
por: Zhao, Yang, et al.
Publicado: (2026)
por: Zhao, Yang, et al.
Publicado: (2026)
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
por: Kim, Gyuhak, et al.
Publicado: (2025)
por: Kim, Gyuhak, et al.
Publicado: (2025)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
por: Liu, Zhihan, et al.
Publicado: (2024)
por: Liu, Zhihan, et al.
Publicado: (2024)
torchtune: PyTorch native post-training library
por: Obozov, Mark, et al.
Publicado: (2026)
por: Obozov, Mark, et al.
Publicado: (2026)
Delayed homomorphic reinforcement learning for environments with delayed feedback
por: Lee, Jongsoo, et al.
Publicado: (2026)
por: Lee, Jongsoo, et al.
Publicado: (2026)
Counterfactual experience augmented off-policy reinforcement learning
por: Lee, Sunbowen, et al.
Publicado: (2025)
por: Lee, Sunbowen, et al.
Publicado: (2025)
Causal prompting model-based offline reinforcement learning
por: Yu, Xuehui, et al.
Publicado: (2024)
por: Yu, Xuehui, et al.
Publicado: (2024)
Deep reinforcement learning with time-scale invariant memory
por: Kabir, Md Rysul, et al.
Publicado: (2024)
por: Kabir, Md Rysul, et al.
Publicado: (2024)
Offline reinforcement learning for job-shop scheduling problems
por: Echeverria, Imanol, et al.
Publicado: (2024)
por: Echeverria, Imanol, et al.
Publicado: (2024)
Bellman operator convergence enhancements in reinforcement learning algorithms
por: Kadurha, David Krame, et al.
Publicado: (2025)
por: Kadurha, David Krame, et al.
Publicado: (2025)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
por: Liu, Zihan, et al.
Publicado: (2025)
por: Liu, Zihan, et al.
Publicado: (2025)
Soup to go: mitigating forgetting during continual learning with model averaging
por: Kleiman, Anat, et al.
Publicado: (2025)
por: Kleiman, Anat, et al.
Publicado: (2025)
Not all tokens are needed(NAT): token efficient reinforcement learning
por: Sang, Hejian, et al.
Publicado: (2026)
por: Sang, Hejian, et al.
Publicado: (2026)
Leveraging weights signals -- Predicting and improving generalizability in reinforcement learning
por: Moulin, Olivier, et al.
Publicado: (2025)
por: Moulin, Olivier, et al.
Publicado: (2025)
Dynamic feature selection in medical predictive monitoring by reinforcement learning
por: Chen, Yutong, et al.
Publicado: (2024)
por: Chen, Yutong, et al.
Publicado: (2024)
Economic span selection of bridge based on deep reinforcement learning
por: Zhang, Leye, et al.
Publicado: (2024)
por: Zhang, Leye, et al.
Publicado: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
por: Berglund, Lukas, et al.
Publicado: (2023)
por: Berglund, Lukas, et al.
Publicado: (2023)
Task diversity produces systematic transfer but inhibits continual reinforcement learning
por: Seth, Purab, et al.
Publicado: (2026)
por: Seth, Purab, et al.
Publicado: (2026)
Found-RL: foundation model-enhanced reinforcement learning for autonomous driving
por: Qu, Yansong, et al.
Publicado: (2026)
por: Qu, Yansong, et al.
Publicado: (2026)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
por: Wu, Xinquan, et al.
Publicado: (2025)
por: Wu, Xinquan, et al.
Publicado: (2025)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
por: Saghafian, Armin, et al.
Publicado: (2024)
por: Saghafian, Armin, et al.
Publicado: (2024)
Ejemplares similares
-
Where does output diversity collapse in post-training?
por: Karouzos, Constantinos, et al.
Publicado: (2026) -
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
por: Kanai, Sekitoshi, et al.
Publicado: (2025) -
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
por: Koh, Woosung, et al.
Publicado: (2026) -
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
por: Peng, Ruiying, et al.
Publicado: (2026) -
Debunk the Myth of SFT Generalization
por: Lin, Xiaofeng, et al.
Publicado: (2025)