A Survey on LLM Mid-Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908627940409344 |
|---|---|
| author | Tu, Chengying Zhang, Xuemiao Weng, Rongxiang Li, Rumei Zhang, Chen Bai, Yang Yan, Hongfei Wang, Jingang Cai, Xunliang |
| author_facet | Tu, Chengying Zhang, Xuemiao Weng, Rongxiang Li, Rumei Zhang, Chen Bai, Yang Yan, Hongfei Wang, Jingang Cai, Xunliang |
| contents | Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23081 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Survey on LLM Mid-Training Tu, Chengying Zhang, Xuemiao Weng, Rongxiang Li, Rumei Zhang, Chen Bai, Yang Yan, Hongfei Wang, Jingang Cai, Xunliang Computation and Language Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs. |
| title | A Survey on LLM Mid-Training |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.23081 |