A Survey on LLM Mid-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Chengying, Zhang, Xuemiao, Weng, Rongxiang, Li, Rumei, Zhang, Chen, Bai, Yang, Yan, Hongfei, Wang, Jingang, Cai, Xunliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908627940409344
author Tu, Chengying
Zhang, Xuemiao
Weng, Rongxiang
Li, Rumei
Zhang, Chen
Bai, Yang
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
author_facet Tu, Chengying
Zhang, Xuemiao
Weng, Rongxiang
Li, Rumei
Zhang, Chen
Bai, Yang
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
contents Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on LLM Mid-Training
Tu, Chengying
Zhang, Xuemiao
Weng, Rongxiang
Li, Rumei
Zhang, Chen
Bai, Yang
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
Computation and Language
Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs.
title A Survey on LLM Mid-Training
topic Computation and Language
url https://arxiv.org/abs/2510.23081