BTS: Harmonizing Specialized Experts into a Generalist LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qizhen, Bhargava, Prajjwal, Bi, Chloe, Cai, Chris X., Foerster, Jakob, Fu, Jeremy, Koura, Punit Singh, Silva, Ruan, Shen, Sheng, Dinan, Emily, Gururangan, Suchin, Lewis, Mike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913673744744448
author Zhang, Qizhen
Bhargava, Prajjwal
Bi, Chloe
Cai, Chris X.
Foerster, Jakob
Fu, Jeremy
Koura, Punit Singh
Silva, Ruan
Shen, Sheng
Dinan, Emily
Gururangan, Suchin
Lewis, Mike
author_facet Zhang, Qizhen
Bhargava, Prajjwal
Bi, Chloe
Cai, Chris X.
Foerster, Jakob
Fu, Jeremy
Koura, Punit Singh
Silva, Ruan
Shen, Sheng
Dinan, Emily
Gururangan, Suchin
Lewis, Mike
contents We present Branch-Train-Stitch (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable generalist model. Following Li et al., we start with a single seed language model which is branched into domain-specific (e.g., coding or math) experts with continual pretraining. BTS combines experts into a generalist model using lightweight stitch layers, which are inserted between frozen experts and the seed LLM, and trained on a small datamix of the expert domains. Stitch layers enable the seed LLM to integrate representations from any number of experts during the forward pass, allowing it to generalize to new domains, despite remaining frozen. Because BTS does not alter the constituent LLMs, BTS provides a modular and flexible approach: experts can be easily removed and new experts can be added with only a small amount of training. Compared to alternative model merging approaches, BTS yields the best generalist performance on a variety of downstream tasks, retaining the specialized capabilities of each of the experts.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BTS: Harmonizing Specialized Experts into a Generalist LLM
Zhang, Qizhen
Bhargava, Prajjwal
Bi, Chloe
Cai, Chris X.
Foerster, Jakob
Fu, Jeremy
Koura, Punit Singh
Silva, Ruan
Shen, Sheng
Dinan, Emily
Gururangan, Suchin
Lewis, Mike
Computation and Language
Machine Learning
We present Branch-Train-Stitch (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable generalist model. Following Li et al., we start with a single seed language model which is branched into domain-specific (e.g., coding or math) experts with continual pretraining. BTS combines experts into a generalist model using lightweight stitch layers, which are inserted between frozen experts and the seed LLM, and trained on a small datamix of the expert domains. Stitch layers enable the seed LLM to integrate representations from any number of experts during the forward pass, allowing it to generalize to new domains, despite remaining frozen. Because BTS does not alter the constituent LLMs, BTS provides a modular and flexible approach: experts can be easily removed and new experts can be added with only a small amount of training. Compared to alternative model merging approaches, BTS yields the best generalist performance on a variety of downstream tasks, retaining the specialized capabilities of each of the experts.
title BTS: Harmonizing Specialized Experts into a Generalist LLM
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.00075