Parallel Scaling Law for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Mouxiang, Hui, Binyuan, Cui, Zeyu, Yang, Jiaxi, Liu, Dayiheng, Sun, Jianling, Lin, Junyang, Liu, Zhongxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916738267873280
author Chen, Mouxiang
Hui, Binyuan
Cui, Zeyu
Yang, Jiaxi
Liu, Dayiheng
Sun, Jianling
Lin, Junyang
Liu, Zhongxin
author_facet Chen, Mouxiang
Hui, Binyuan
Cui, Zeyu
Yang, Jiaxi
Liu, Dayiheng
Sun, Jianling
Lin, Junyang
Liu, Zhongxin
contents It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce the third and more inference-efficient scaling paradigm: increasing the model's parallel computation during both training and inference time. We apply $P$ diverse and learnable transformations to the input, execute forward passes of the model in parallel, and dynamically aggregate the $P$ outputs. This method, namely parallel scaling (ParScale), scales parallel computation by reusing existing parameters and can be applied to any model structure, optimization procedure, data, or task. We theoretically propose a new scaling law and validate it through large-scale pre-training, which shows that a model with $P$ parallel streams is similar to scaling the parameters by $O(\log P)$ while showing superior inference efficiency. For example, ParScale can use up to 22$\times$ less memory increase and 6$\times$ less latency increase compared to parameter scaling that achieves the same performance improvement. It can also recycle an off-the-shelf pre-trained model into a parallelly scaled one by post-training on a small amount of tokens, further reducing the training budget. The new scaling law we discovered potentially facilitates the deployment of more powerful models in low-resource scenarios, and provides an alternative perspective for the role of computation in machine learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parallel Scaling Law for Language Models
Chen, Mouxiang
Hui, Binyuan
Cui, Zeyu
Yang, Jiaxi
Liu, Dayiheng
Sun, Jianling
Lin, Junyang
Liu, Zhongxin
Machine Learning
Computation and Language
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce the third and more inference-efficient scaling paradigm: increasing the model's parallel computation during both training and inference time. We apply $P$ diverse and learnable transformations to the input, execute forward passes of the model in parallel, and dynamically aggregate the $P$ outputs. This method, namely parallel scaling (ParScale), scales parallel computation by reusing existing parameters and can be applied to any model structure, optimization procedure, data, or task. We theoretically propose a new scaling law and validate it through large-scale pre-training, which shows that a model with $P$ parallel streams is similar to scaling the parameters by $O(\log P)$ while showing superior inference efficiency. For example, ParScale can use up to 22$\times$ less memory increase and 6$\times$ less latency increase compared to parameter scaling that achieves the same performance improvement. It can also recycle an off-the-shelf pre-trained model into a parallelly scaled one by post-training on a small amount of tokens, further reducing the training budget. The new scaling law we discovered potentially facilitates the deployment of more powerful models in low-resource scenarios, and provides an alternative perspective for the role of computation in machine learning.
title Parallel Scaling Law for Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.10475