Optimization Hyper-parameter Laws for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Xingyu, Ding, Kuangyu, Yan, Shuicheng, Toh, Kim-Chuan, Wei, Tianwen
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916029392748544
author Xie, Xingyu
Ding, Kuangyu
Yan, Shuicheng
Toh, Kim-Chuan
Wei, Tianwen
author_facet Xie, Xingyu
Ding, Kuangyu
Yan, Shuicheng
Toh, Kim-Chuan
Wei, Tianwen
contents Large Language Models have driven significant AI advancements, yet their training is resource-intensive and highly sensitive to hyper-parameter selection. While scaling laws provide valuable guidance on model size and data requirements, they fall short in choosing dynamic hyper-parameters, such as learning-rate (LR) schedules, that evolve during training. To bridge this gap, we present Optimization Hyper-parameter Laws (Opt-Laws), a framework that predicts final training loss as a function of LR schedule, model size, and data size. Grounded in SDE-based convergence and escape analyses, Opt-Laws yield interpretable convergence and escape features that predict final training loss across model scales, enabling schedule pre-selection from small-scale experiments. Empirically, Opt-Laws achieve a 94% Top-2 hit rate for identifying near-optimal schedule candidates on held-out configurations, correctly identify the best-performing schedule family in all five evaluated out-of-family settings, and detect training divergence with F1 = 0.92.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04777
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimization Hyper-parameter Laws for Large Language Models
Xie, Xingyu
Ding, Kuangyu
Yan, Shuicheng
Toh, Kim-Chuan
Wei, Tianwen
Machine Learning
Optimization and Control
Large Language Models have driven significant AI advancements, yet their training is resource-intensive and highly sensitive to hyper-parameter selection. While scaling laws provide valuable guidance on model size and data requirements, they fall short in choosing dynamic hyper-parameters, such as learning-rate (LR) schedules, that evolve during training. To bridge this gap, we present Optimization Hyper-parameter Laws (Opt-Laws), a framework that predicts final training loss as a function of LR schedule, model size, and data size. Grounded in SDE-based convergence and escape analyses, Opt-Laws yield interpretable convergence and escape features that predict final training loss across model scales, enabling schedule pre-selection from small-scale experiments. Empirically, Opt-Laws achieve a 94% Top-2 hit rate for identifying near-optimal schedule candidates on held-out configurations, correctly identify the best-performing schedule family in all five evaluated out-of-family settings, and detect training divergence with F1 = 0.92.
title Optimization Hyper-parameter Laws for Large Language Models
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2409.04777