mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
Fuente:
arXiv
Salvato in:
| Autori principali: | Koh, Woosung, Jeon, Jeyoung, Song, Youngjin, Cheon, Yujin, Oh, Soowon, Choi, Jaehyeong, Yun, Se-Young |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025)
di: Song, Junha, et al.
Pubblicazione: (2025)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
di: Shin, Haebin, et al.
Pubblicazione: (2025)
di: Shin, Haebin, et al.
Pubblicazione: (2025)
Debunk the Myth of SFT Generalization
di: Lin, Xiaofeng, et al.
Pubblicazione: (2025)
di: Lin, Xiaofeng, et al.
Pubblicazione: (2025)
Predicting LLM Reasoning Performance with Small Proxy Model
di: Koh, Woosung, et al.
Pubblicazione: (2025)
di: Koh, Woosung, et al.
Pubblicazione: (2025)
Generative Visual Code Mobile World Models
di: Koh, Woosung, et al.
Pubblicazione: (2026)
di: Koh, Woosung, et al.
Pubblicazione: (2026)
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
di: Koh, Woosung, et al.
Pubblicazione: (2024)
di: Koh, Woosung, et al.
Pubblicazione: (2024)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
di: Oh, Jihwan, et al.
Pubblicazione: (2026)
di: Oh, Jihwan, et al.
Pubblicazione: (2026)
Crowd-SFT: Crowdsourcing for LLM Alignment
di: Sotiropoulos, Alex, et al.
Pubblicazione: (2025)
di: Sotiropoulos, Alex, et al.
Pubblicazione: (2025)
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
di: Li, Jiaxiang, et al.
Pubblicazione: (2024)
di: Li, Jiaxiang, et al.
Pubblicazione: (2024)
A Three-Dimensional SFT with Sparse Columns
di: Salo, Ville, et al.
Pubblicazione: (2025)
di: Salo, Ville, et al.
Pubblicazione: (2025)
Simplified SFT moduli spaces for Legendrian links
di: Avdek, Russell
Pubblicazione: (2021)
di: Avdek, Russell
Pubblicazione: (2021)
SFT covers for actions of the first Grigorchuk group
di: Grigorchuk, Rostislav, et al.
Pubblicazione: (2024)
di: Grigorchuk, Rostislav, et al.
Pubblicazione: (2024)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
di: Koh, Woosung, et al.
Pubblicazione: (2024)
di: Koh, Woosung, et al.
Pubblicazione: (2024)
A landscape of contact manifolds via rational SFT
di: Moreno, Agustin, et al.
Pubblicazione: (2020)
di: Moreno, Agustin, et al.
Pubblicazione: (2020)
RL Fine-Tuning Heals OOD Forgetting in SFT
di: Jin, Hangzhan, et al.
Pubblicazione: (2025)
di: Jin, Hangzhan, et al.
Pubblicazione: (2025)
SFT for ASD: A systemic intervention for neurodiverse families
di: Anthony Pennant
Pubblicazione: (2024)
di: Anthony Pennant
Pubblicazione: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
Crafting Reversible SFT Behaviors in Large Language Models
di: Lin, Yuping, et al.
Pubblicazione: (2026)
di: Lin, Yuping, et al.
Pubblicazione: (2026)
Learning to Adapt SFT Data for Better Reasoning Generalization
di: Sun, Lisong, et al.
Pubblicazione: (2026)
di: Sun, Lisong, et al.
Pubblicazione: (2026)
On Countable SFT Covers of Sparse Multidimensional Shift Spaces
di: Törmä, Ilkka
Pubblicazione: (2024)
di: Törmä, Ilkka
Pubblicazione: (2024)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
SED-SFT: Selectively Encouraging Diversity in Supervised Fine-Tuning
di: Chen, Yijie, et al.
Pubblicazione: (2026)
di: Chen, Yijie, et al.
Pubblicazione: (2026)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
di: Ou, Linyu, et al.
Pubblicazione: (2025)
di: Ou, Linyu, et al.
Pubblicazione: (2025)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
di: Zhu, Taojie, et al.
Pubblicazione: (2026)
di: Zhu, Taojie, et al.
Pubblicazione: (2026)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective
di: Zhang, Junpeng, et al.
Pubblicazione: (2026)
di: Zhang, Junpeng, et al.
Pubblicazione: (2026)
TMS: Trajectory-Mixed Supervision for Reward-Free, On-Policy SFT
di: Khan, Rana Muhammad Shahroz, et al.
Pubblicazione: (2026)
di: Khan, Rana Muhammad Shahroz, et al.
Pubblicazione: (2026)
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
di: Wu, Yongliang, et al.
Pubblicazione: (2025)
di: Wu, Yongliang, et al.
Pubblicazione: (2025)
Comments on resolution of nonassociativity in SFT- an example from axioms of BCFT-
di: Matsuo Yutaka
Pubblicazione: (2002)
di: Matsuo Yutaka
Pubblicazione: (2002)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
di: Peng, Ruiying, et al.
Pubblicazione: (2026)
di: Peng, Ruiying, et al.
Pubblicazione: (2026)
Gradients Must Earn Their Influence: Unifying SFT with Generalized Entropic Objectives
di: Wang, Zecheng, et al.
Pubblicazione: (2026)
di: Wang, Zecheng, et al.
Pubblicazione: (2026)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
di: Feng, Yuming, et al.
Pubblicazione: (2026)
di: Feng, Yuming, et al.
Pubblicazione: (2026)
RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following
di: Wang, Zhichao, et al.
Pubblicazione: (2025)
di: Wang, Zhichao, et al.
Pubblicazione: (2025)
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization
di: Su, Xiaole, et al.
Pubblicazione: (2026)
di: Su, Xiaole, et al.
Pubblicazione: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?
di: Gu, Weizheng, et al.
Pubblicazione: (2026)
di: Gu, Weizheng, et al.
Pubblicazione: (2026)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025) -
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
di: Shin, Haebin, et al.
Pubblicazione: (2025) -
Debunk the Myth of SFT Generalization
di: Lin, Xiaofeng, et al.
Pubblicazione: (2025) -
Predicting LLM Reasoning Performance with Small Proxy Model
di: Koh, Woosung, et al.
Pubblicazione: (2025) -
Generative Visual Code Mobile World Models
di: Koh, Woosung, et al.
Pubblicazione: (2026)