Shadow-FT: Tuning Instruct Model via Training on Paired Base Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Taiqiang, Yang, Runming, Li, Jiayi, Hu, Pengfei, Wu, Yik-Chung, Wong, Ngai, Yang, Yujiu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911177405104128
author Wu, Taiqiang
Yang, Runming
Li, Jiayi
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
author_facet Wu, Taiqiang
Yang, Runming
Li, Jiayi
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
contents Large language models (LLMs) consistently benefit from further fine-tuning on various tasks. However, we observe that directly tuning the Instruct (i.e., instruction-tuned) models often leads to marginal improvements and even performance degeneration. Notably, paired Base models, the foundation for these Instruct variants, contain highly similar weight values (i.e., less than 2% on average for Llama 3.1 8B). The Base model tends to be a good learner yet a weak backbone without post-training. Therefore, we propose a novel Shadow-FT framework to tune the Instruct models by leveraging the corresponding Base models. The key insight is to fine-tune the Base model, and then \textit{directly} graft the learned weight updates to the Instruct model. Our proposed Shadow-FT introduces no additional parameters, is easy to implement, and significantly improves performance. We conduct extensive experiments on tuning mainstream LLMs, such as Qwen 3 and Llama 3 series, and evaluate them across 19 benchmarks covering coding, reasoning, and mathematical tasks. Experimental results demonstrate that Shadow-FT consistently outperforms conventional full-parameter and parameter-efficient tuning approaches. Further analyses indicate that Shadow-FT can be applied to multimodal large language models (MLLMs) and combined with direct preference optimization~(DPO). Codes and weights are available at \href{https://github.com/wutaiqiang/Shadow-FT}{Github}.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12716
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
Wu, Taiqiang
Yang, Runming
Li, Jiayi
Hu, Pengfei
Wu, Yik-Chung
Wong, Ngai
Yang, Yujiu
Computation and Language
Artificial Intelligence
Large language models (LLMs) consistently benefit from further fine-tuning on various tasks. However, we observe that directly tuning the Instruct (i.e., instruction-tuned) models often leads to marginal improvements and even performance degeneration. Notably, paired Base models, the foundation for these Instruct variants, contain highly similar weight values (i.e., less than 2% on average for Llama 3.1 8B). The Base model tends to be a good learner yet a weak backbone without post-training. Therefore, we propose a novel Shadow-FT framework to tune the Instruct models by leveraging the corresponding Base models. The key insight is to fine-tune the Base model, and then \textit{directly} graft the learned weight updates to the Instruct model. Our proposed Shadow-FT introduces no additional parameters, is easy to implement, and significantly improves performance. We conduct extensive experiments on tuning mainstream LLMs, such as Qwen 3 and Llama 3 series, and evaluate them across 19 benchmarks covering coding, reasoning, and mathematical tasks. Experimental results demonstrate that Shadow-FT consistently outperforms conventional full-parameter and parameter-efficient tuning approaches. Further analyses indicate that Shadow-FT can be applied to multimodal large language models (MLLMs) and combined with direct preference optimization~(DPO). Codes and weights are available at \href{https://github.com/wutaiqiang/Shadow-FT}{Github}.
title Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.12716