DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Haebin, Ji, Lei, Liu, Xiao, Yu, Zhiwei, Chen, Qi, Gong, Yeyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
by: Shin, Haebin, et al.
Published: (2025)
by: Shin, Haebin, et al.
Published: (2025)
Generative Prompt Internalization
by: Shin, Haebin, et al.
Published: (2024)
by: Shin, Haebin, et al.
Published: (2024)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025)
by: Yang, Kailai, et al.
Published: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
by: Zhu, Jingwei, et al.
Published: (2024)
by: Zhu, Jingwei, et al.
Published: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing
by: Duan, Yiwen, et al.
Published: (2024)
by: Duan, Yiwen, et al.
Published: (2024)
Contrastive Instruction Tuning
by: Yan, Tianyi Lorena, et al.
Published: (2024)
by: Yan, Tianyi Lorena, et al.
Published: (2024)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Generative Representational Instruction Tuning
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
by: Du, Yuhao, et al.
Published: (2025)
by: Du, Yuhao, et al.
Published: (2025)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
by: Yuan, Zhihang, et al.
Published: (2026)
by: Yuan, Zhihang, et al.
Published: (2026)
Gradients Must Earn Their Influence: Unifying SFT with Generalized Entropic Objectives
by: Wang, Zecheng, et al.
Published: (2026)
by: Wang, Zecheng, et al.
Published: (2026)
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)
by: Köksal, Abdullatif, et al.
Published: (2023)
Instruction Tuning with Human Curriculum
by: Lee, Bruce W., et al.
Published: (2023)
by: Lee, Bruce W., et al.
Published: (2023)
Continual SFT Matches Multimodal RLHF with Negative Supervision
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Enhancing Large Language Model Performance with Gradient-Based Parameter Selection
by: Li, Haoling, et al.
Published: (2024)
by: Li, Haoling, et al.
Published: (2024)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
Integrative Decoding: Improve Factuality via Implicit Self-consistency
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
On the Effect of Instruction Tuning Loss on Generalization
by: Chatterjee, Anwoy, et al.
Published: (2025)
by: Chatterjee, Anwoy, et al.
Published: (2025)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
by: Zeng, Runjia, et al.
Published: (2025)
by: Zeng, Runjia, et al.
Published: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
DynaPrompt: Dynamic Test-Time Prompt Tuning
by: Xiao, Zehao, et al.
Published: (2025)
by: Xiao, Zehao, et al.
Published: (2025)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
The Best Instruction-Tuning Data are Those That Fit
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Parameter Efficient Instruction Tuning: An Empirical Study
by: He, Pengfei
Published: (2024)
by: He, Pengfei
Published: (2024)
Similar Items
-
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
by: Shin, Haebin, et al.
Published: (2025) -
Generative Prompt Internalization
by: Shin, Haebin, et al.
Published: (2024) -
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025) -
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024) -
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)