ZClip: Adaptive Spike Mitigation for LLM Pre-Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kumar, Abhay, Owen, Louis, Chowdhury, Nilabhra Roy, Güra, Fabian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917975730159616
author Kumar, Abhay
Owen, Louis
Chowdhury, Nilabhra Roy
Güra, Fabian
author_facet Kumar, Abhay
Owen, Louis
Chowdhury, Nilabhra Roy
Güra, Fabian
contents Training large language models (LLMs) presents numerous challenges, including gradient instability and loss spikes. These phenomena can lead to catastrophic divergence, requiring costly checkpoint restoration and data batch skipping. Traditional gradient clipping techniques, such as constant or norm-based methods, fail to address these issues effectively due to their reliance on fixed thresholds or heuristics, leading to inefficient learning and requiring frequent manual intervention. In this work, we propose ZClip, an adaptive gradient clipping algorithm that dynamically adjusts the clipping threshold based on statistical properties of gradient norms over time. Unlike prior reactive strategies, ZClip proactively adapts to training dynamics without making any prior assumptions on the scale and the temporal evolution of gradient norms. At its core, it leverages z-score-based anomaly detection to identify and mitigate large gradient spikes, preventing malignant loss spikes while not interfering with convergence otherwise. Our code is available at: https://github.com/bluorion-com/ZClip.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZClip: Adaptive Spike Mitigation for LLM Pre-Training
Kumar, Abhay
Owen, Louis
Chowdhury, Nilabhra Roy
Güra, Fabian
Machine Learning
Computation and Language
Training large language models (LLMs) presents numerous challenges, including gradient instability and loss spikes. These phenomena can lead to catastrophic divergence, requiring costly checkpoint restoration and data batch skipping. Traditional gradient clipping techniques, such as constant or norm-based methods, fail to address these issues effectively due to their reliance on fixed thresholds or heuristics, leading to inefficient learning and requiring frequent manual intervention. In this work, we propose ZClip, an adaptive gradient clipping algorithm that dynamically adjusts the clipping threshold based on statistical properties of gradient norms over time. Unlike prior reactive strategies, ZClip proactively adapts to training dynamics without making any prior assumptions on the scale and the temporal evolution of gradient norms. At its core, it leverages z-score-based anomaly detection to identify and mitigate large gradient spikes, preventing malignant loss spikes while not interfering with convergence otherwise. Our code is available at: https://github.com/bluorion-com/ZClip.
title ZClip: Adaptive Spike Mitigation for LLM Pre-Training
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2504.02507