Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shuhe, Wang, Guoyin, Wang, Yizhong, Li, Jiwei, Hovy, Eduard, Guo, Chen
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915006851842048
author Wang, Shuhe
Wang, Guoyin
Wang, Yizhong
Li, Jiwei
Hovy, Eduard
Guo, Chen
author_facet Wang, Shuhe
Wang, Guoyin
Wang, Yizhong
Li, Jiwei
Hovy, Eduard
Guo, Chen
contents Packing, initially utilized in the pre-training phase, is an optimization technique designed to maximize hardware resource efficiency by combining different training sequences to fit the model's maximum input length. Although it has demonstrated effectiveness during pre-training, there remains a lack of comprehensive analysis for the supervised fine-tuning (SFT) stage on the following points: (1) whether packing can effectively enhance training efficiency while maintaining performance, (2) the suitable size of the model and dataset for fine-tuning with the packing method, and (3) whether packing unrelated or related training samples might cause the model to either excessively disregard or over-rely on the context. In this paper, we perform extensive comparisons between SFT methods using padding and packing, covering SFT datasets ranging from 69K to 1.2M and models from 8B to 70B. This provides the first comprehensive analysis of the advantages and limitations of packing versus padding, as well as practical considerations for implementing packing in various training scenarios. Our analysis covers various benchmarks, including knowledge, reasoning, and coding, as well as GPT-based evaluations, time efficiency, and other fine-tuning parameters. We also open-source our code for fine-tuning and evaluation and provide checkpoints fine-tuned on datasets of different sizes, aiming to advance future research on packing methods. Code is available at: https://github.com/ShuheWang1998/Packing-Analysis?tab=readme-ov-file.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
Wang, Shuhe
Wang, Guoyin
Wang, Yizhong
Li, Jiwei
Hovy, Eduard
Guo, Chen
Machine Learning
Artificial Intelligence
Computation and Language
Packing, initially utilized in the pre-training phase, is an optimization technique designed to maximize hardware resource efficiency by combining different training sequences to fit the model's maximum input length. Although it has demonstrated effectiveness during pre-training, there remains a lack of comprehensive analysis for the supervised fine-tuning (SFT) stage on the following points: (1) whether packing can effectively enhance training efficiency while maintaining performance, (2) the suitable size of the model and dataset for fine-tuning with the packing method, and (3) whether packing unrelated or related training samples might cause the model to either excessively disregard or over-rely on the context. In this paper, we perform extensive comparisons between SFT methods using padding and packing, covering SFT datasets ranging from 69K to 1.2M and models from 8B to 70B. This provides the first comprehensive analysis of the advantages and limitations of packing versus padding, as well as practical considerations for implementing packing in various training scenarios. Our analysis covers various benchmarks, including knowledge, reasoning, and coding, as well as GPT-based evaluations, time efficiency, and other fine-tuning parameters. We also open-source our code for fine-tuning and evaluation and provide checkpoints fine-tuned on datasets of different sizes, aiming to advance future research on packing methods. Code is available at: https://github.com/ShuheWang1998/Packing-Analysis?tab=readme-ov-file.
title Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.08081