Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Hong, Li, Zhiyuan, Hall, David, Liang, Percy, Ma, Tengyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917604696784896
author Liu, Hong
Li, Zhiyuan
Hall, David
Liang, Percy
Ma, Tengyu
author_facet Liu, Hong
Li, Zhiyuan
Hall, David
Liang, Percy
Ma, Tengyu
contents Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimizers often incur too much per-step overhead. In this paper, we propose Sophia, Second-order Clipped Stochastic Optimization, a simple scalable second-order optimizer that uses a light-weight estimate of the diagonal Hessian as the pre-conditioner. The update is the moving average of the gradients divided by the moving average of the estimated Hessian, followed by element-wise clipping. The clipping controls the worst-case update size and tames the negative impact of non-convexity and rapid change of Hessian along the trajectory. Sophia only estimates the diagonal Hessian every handful of iterations, which has negligible average per-step time and memory overhead. On language modeling with GPT models of sizes ranging from 125M to 1.5B, Sophia achieves a 2x speed-up compared to Adam in the number of steps, total compute, and wall-clock time, achieving the same perplexity with 50% fewer steps, less total compute, and reduced wall-clock time. Theoretically, we show that Sophia, in a much simplified setting, adapts to the heterogeneous curvatures in different parameter dimensions, and thus has a run-time bound that does not depend on the condition number of the loss.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14342
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Liu, Hong
Li, Zhiyuan
Hall, David
Liang, Percy
Ma, Tengyu
Machine Learning
Computation and Language
Optimization and Control
Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimizers often incur too much per-step overhead. In this paper, we propose Sophia, Second-order Clipped Stochastic Optimization, a simple scalable second-order optimizer that uses a light-weight estimate of the diagonal Hessian as the pre-conditioner. The update is the moving average of the gradients divided by the moving average of the estimated Hessian, followed by element-wise clipping. The clipping controls the worst-case update size and tames the negative impact of non-convexity and rapid change of Hessian along the trajectory. Sophia only estimates the diagonal Hessian every handful of iterations, which has negligible average per-step time and memory overhead. On language modeling with GPT models of sizes ranging from 125M to 1.5B, Sophia achieves a 2x speed-up compared to Adam in the number of steps, total compute, and wall-clock time, achieving the same perplexity with 50% fewer steps, less total compute, and reduced wall-clock time. Theoretically, we show that Sophia, in a much simplified setting, adapts to the heterogeneous curvatures in different parameter dimensions, and thus has a run-time bound that does not depend on the condition number of the loss.
title Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
topic Machine Learning
Computation and Language
Optimization and Control
url https://arxiv.org/abs/2305.14342