Llama 3 Meets MoE: Efficient Upcycling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vavre, Aditya, He, Ethan, Liu, Dennis, Yan, Zijie, Yang, June, Tajbakhsh, Nima, Aithal, Ashwath
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910743840948224
author Vavre, Aditya
He, Ethan
Liu, Dennis
Yan, Zijie
Yang, June
Tajbakhsh, Nima
Aithal, Ashwath
author_facet Vavre, Aditya
He, Ethan
Liu, Dennis
Yan, Zijie
Yang, June
Tajbakhsh, Nima
Aithal, Ashwath
contents Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternative, increasing capacity without a proportional rise in compute requirements. However, training MoE models from scratch poses challenges like overfitting and routing instability. We present an efficient training recipe leveraging pre-trained dense checkpoints, training an 8-Expert Top-2 MoE model from Llama 3-8B with less than $1\%$ of typical pre-training compute. Our approach enhances downstream performance on academic benchmarks, achieving a $\textbf{2%}$ improvement in 0-shot accuracy on MMLU, while reaching a Model FLOPs Utilization (MFU) of $\textbf{46.8%}$ during training using our framework. We also integrate online upcycling in NeMo for seamless use of pre-trained weights, enabling cost-effective development of high-capacity MoE models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09952
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Llama 3 Meets MoE: Efficient Upcycling
Vavre, Aditya
He, Ethan
Liu, Dennis
Yan, Zijie
Yang, June
Tajbakhsh, Nima
Aithal, Ashwath
Machine Learning
Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternative, increasing capacity without a proportional rise in compute requirements. However, training MoE models from scratch poses challenges like overfitting and routing instability. We present an efficient training recipe leveraging pre-trained dense checkpoints, training an 8-Expert Top-2 MoE model from Llama 3-8B with less than $1\%$ of typical pre-training compute. Our approach enhances downstream performance on academic benchmarks, achieving a $\textbf{2%}$ improvement in 0-shot accuracy on MMLU, while reaching a Model FLOPs Utilization (MFU) of $\textbf{46.8%}$ during training using our framework. We also integrate online upcycling in NeMo for seamless use of pre-trained weights, enabling cost-effective development of high-capacity MoE models.
title Llama 3 Meets MoE: Efficient Upcycling
topic Machine Learning
url https://arxiv.org/abs/2412.09952