Improving Continual Pre-training Through Seamless Data Packing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yin, Ruicheng, Gao, Xuan, Lv, Changze, Wang, Xiaohua, Zheng, Xiaoqing, Huang, Xuanjing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916766040457216
author Yin, Ruicheng
Gao, Xuan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Yin, Ruicheng
Gao, Xuan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
contents Continual pre-training has demonstrated significant potential in enhancing model performance, particularly in domain-specific scenarios. The most common approach for packing data before continual pre-training involves concatenating input texts and splitting them into fixed-length sequences. While straightforward and efficient, this method often leads to excessive truncation and context discontinuity, which can hinder model performance. To address these issues, we explore the potential of data engineering to enhance continual pre-training, particularly its impact on model performance and efficiency. We propose Seamless Packing (SP), a novel data packing strategy aimed at preserving contextual information more effectively and enhancing model performance. Our approach employs a sliding window technique in the first stage that synchronizes overlapping tokens across consecutive sequences, ensuring better continuity and contextual coherence. In the second stage, we adopt a First-Fit-Decreasing algorithm to pack shorter texts into bins slightly larger than the target sequence length, thereby minimizing padding and truncation. Empirical evaluations across various model architectures and corpus domains demonstrate the effectiveness of our method, outperforming baseline method in 99% of all settings. Code is available at https://github.com/Infernus-WIND/Seamless-Packing.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22018
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Continual Pre-training Through Seamless Data Packing
Yin, Ruicheng
Gao, Xuan
Lv, Changze
Wang, Xiaohua
Zheng, Xiaoqing
Huang, Xuanjing
Computation and Language
Continual pre-training has demonstrated significant potential in enhancing model performance, particularly in domain-specific scenarios. The most common approach for packing data before continual pre-training involves concatenating input texts and splitting them into fixed-length sequences. While straightforward and efficient, this method often leads to excessive truncation and context discontinuity, which can hinder model performance. To address these issues, we explore the potential of data engineering to enhance continual pre-training, particularly its impact on model performance and efficiency. We propose Seamless Packing (SP), a novel data packing strategy aimed at preserving contextual information more effectively and enhancing model performance. Our approach employs a sliding window technique in the first stage that synchronizes overlapping tokens across consecutive sequences, ensuring better continuity and contextual coherence. In the second stage, we adopt a First-Fit-Decreasing algorithm to pack shorter texts into bins slightly larger than the target sequence length, thereby minimizing padding and truncation. Empirical evaluations across various model architectures and corpus domains demonstrate the effectiveness of our method, outperforming baseline method in 99% of all settings. Code is available at https://github.com/Infernus-WIND/Seamless-Packing.
title Improving Continual Pre-training Through Seamless Data Packing
topic Computation and Language
url https://arxiv.org/abs/2505.22018