Saved in:
Bibliographic Details
Main Authors: Liew, Seng Pei, Kato, Takuya, Takase, Sho
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.03009
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918059738923008
author Liew, Seng Pei
Kato, Takuya
Takase, Sho
author_facet Liew, Seng Pei
Kato, Takuya
Takase, Sho
contents Pretraining large language models (LLMs) is resource-intensive, often requiring months of training time even with high-end GPU clusters. There are two approaches of mitigating such computational demands: reusing smaller models to train larger ones (upcycling), and training computationally efficient models like mixture-of-experts (MoE). In this paper, we study the upcycling of LLMs to MoE models, of which the scaling behavior remains underexplored. Through extensive experiments, we identify empirical scaling laws that describe how performance depends on dataset size and model configuration. Particularly, we show that, while scaling these factors improves performance, there is a novel interaction term between the dense and upcycled training dataset that limits the efficiency of upcycling at large computational budgets. Based on these findings, we provide guidance to scale upcycling, and establish conditions under which upcycling outperforms from-scratch trainings within budget constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Laws for Upcycling Mixture-of-Experts Language Models
Liew, Seng Pei
Kato, Takuya
Takase, Sho
Machine Learning
Computation and Language
Pretraining large language models (LLMs) is resource-intensive, often requiring months of training time even with high-end GPU clusters. There are two approaches of mitigating such computational demands: reusing smaller models to train larger ones (upcycling), and training computationally efficient models like mixture-of-experts (MoE). In this paper, we study the upcycling of LLMs to MoE models, of which the scaling behavior remains underexplored. Through extensive experiments, we identify empirical scaling laws that describe how performance depends on dataset size and model configuration. Particularly, we show that, while scaling these factors improves performance, there is a novel interaction term between the dense and upcycled training dataset that limits the efficiency of upcycling at large computational budgets. Based on these findings, we provide guidance to scale upcycling, and establish conditions under which upcycling outperforms from-scratch trainings within budget constraints.
title Scaling Laws for Upcycling Mixture-of-Experts Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2502.03009