Reusing Overtrained Language Models Saturates Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liew, Seng Pei, Kato, Takuya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918320314253312
author Liew, Seng Pei
Kato, Takuya
author_facet Liew, Seng Pei
Kato, Takuya
contents Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. However, the effectiveness remains unclear, especially when applied to overtrained base models. In this work, we empirically study the scaling properties of model reuse and find that the scaling efficiency diminishes in a predictable manner: The scaling exponent with respect to second-stage training tokens decreases logarithmically with the number of tokens used to pretrain the base model. The joint dependence on first- and second-stage tokens is accurately modeled by a simple scaling law. Such saturation effect reveals a fundamental trade-off in multi-stage pretraining strategies: the more extensively a base model is pretrained, the less benefit additional pretraining provides. Our findings provide practical insights for efficient language model training and raise important considerations for the reuse of overtrained models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06548
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reusing Overtrained Language Models Saturates Scaling
Liew, Seng Pei
Kato, Takuya
Computation and Language
Machine Learning
Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. However, the effectiveness remains unclear, especially when applied to overtrained base models. In this work, we empirically study the scaling properties of model reuse and find that the scaling efficiency diminishes in a predictable manner: The scaling exponent with respect to second-stage training tokens decreases logarithmically with the number of tokens used to pretrain the base model. The joint dependence on first- and second-stage tokens is accurately modeled by a simple scaling law. Such saturation effect reveals a fundamental trade-off in multi-stage pretraining strategies: the more extensively a base model is pretrained, the less benefit additional pretraining provides. Our findings provide practical insights for efficient language model training and raise important considerations for the reuse of overtrained models.
title Reusing Overtrained Language Models Saturates Scaling
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.06548