Generalizing Scaling Laws for Dense and Sparse Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hossain, Md Arafat, Wu, Xingfu, Taylor, Valerie, Jannesari, Ali
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914317072334848
author Hossain, Md Arafat
Wu, Xingfu
Taylor, Valerie
Jannesari, Ali
author_facet Hossain, Md Arafat
Wu, Xingfu
Taylor, Valerie
Jannesari, Ali
contents Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Several efforts have addressed the challenge by proposing different empirical scaling laws, but almost all of them are architecture-specific (dense or sparse). In this work we revisit existing empirical scaling laws and propose a generalized scaling law to provide a unified framework that is applicable to both dense and sparse large language models. We evaluate and compare our proposed scaling law with existing scaling laws and demonstrate that our proposed scaling law captures the scaling behavior of existing scaling laws. Further, we show an IsoFLOP comparison between our proposed scaling law and the state-of-the-art scaling law to illustrate the effectiveness of our proposed scaling law for Mixture-of-Expert (MoE)-based very large LLMs like DeepSeek-V3. Our proposed scaling law can be used to estimate the best model hyperparameters (Model size, Tokens and Compute) for a given sparsity or to identify the optimal sparsity for the given model hyperparameters.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06617
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalizing Scaling Laws for Dense and Sparse Large Language Models
Hossain, Md Arafat
Wu, Xingfu
Taylor, Valerie
Jannesari, Ali
Machine Learning
Artificial Intelligence
Performance
Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Several efforts have addressed the challenge by proposing different empirical scaling laws, but almost all of them are architecture-specific (dense or sparse). In this work we revisit existing empirical scaling laws and propose a generalized scaling law to provide a unified framework that is applicable to both dense and sparse large language models. We evaluate and compare our proposed scaling law with existing scaling laws and demonstrate that our proposed scaling law captures the scaling behavior of existing scaling laws. Further, we show an IsoFLOP comparison between our proposed scaling law and the state-of-the-art scaling law to illustrate the effectiveness of our proposed scaling law for Mixture-of-Expert (MoE)-based very large LLMs like DeepSeek-V3. Our proposed scaling law can be used to estimate the best model hyperparameters (Model size, Tokens and Compute) for a given sparsity or to identify the optimal sparsity for the given model hyperparameters.
title Generalizing Scaling Laws for Dense and Sparse Large Language Models
topic Machine Learning
Artificial Intelligence
Performance
url https://arxiv.org/abs/2508.06617