LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Minqiu, Huang, Xin, Guo, Luanzheng, Tallent, Nathan R., Sato, Kento, Dai, Dong
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908852535951360
author Sun, Minqiu
Huang, Xin
Guo, Luanzheng
Tallent, Nathan R.
Sato, Kento
Dai, Dong
author_facet Sun, Minqiu
Huang, Xin
Guo, Luanzheng
Tallent, Nathan R.
Sato, Kento
Dai, Dong
contents Checkpointing is essential for fault tolerance in training large language models (LLMs). However, existing methods, regardless of their I/O strategies, periodically store the entire model and optimizer states, incurring substantial storage overhead and resource contention. Recent studies reveal that updates across LLM layers are highly non-uniform. Across training steps, some layers may undergo more significant changes, while others remain relatively stable or even unchanged. This suggests that selectively checkpointing only layers with significant updates could reduce overhead without harming training. Implementing such selective strategies requires fine-grained control over both weights and optimizer states, which no current tool provides. To address this gap, we propose \texttt{LLMTailor}, a checkpoint-merging framework that filters and assembles layers from different checkpoints to form a composite checkpoint. Our evaluation indicates that LLMTailor can work with different selective checkpointing strategies and effectively reduce checkpoint size (e.g., 4.3 times smaller for Llama3.1-8B) and checkpoint time (e.g., 2.8 times faster for Qwen2.5-7B) while maintaining model quality.
format Preprint
id arxiv_https___arxiv_org_abs_2602_22158
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
Sun, Minqiu
Huang, Xin
Guo, Luanzheng
Tallent, Nathan R.
Sato, Kento
Dai, Dong
Distributed, Parallel, and Cluster Computing
Checkpointing is essential for fault tolerance in training large language models (LLMs). However, existing methods, regardless of their I/O strategies, periodically store the entire model and optimizer states, incurring substantial storage overhead and resource contention. Recent studies reveal that updates across LLM layers are highly non-uniform. Across training steps, some layers may undergo more significant changes, while others remain relatively stable or even unchanged. This suggests that selectively checkpointing only layers with significant updates could reduce overhead without harming training. Implementing such selective strategies requires fine-grained control over both weights and optimizer states, which no current tool provides. To address this gap, we propose \texttt{LLMTailor}, a checkpoint-merging framework that filters and assembles layers from different checkpoints to form a composite checkpoint. Our evaluation indicates that LLMTailor can work with different selective checkpointing strategies and effectively reduce checkpoint size (e.g., 4.3 times smaller for Llama3.1-8B) and checkpoint time (e.g., 2.8 times faster for Qwen2.5-7B) while maintaining model quality.
title LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.22158