Landscape-Aware Growing: The Power of a Little LAG

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karp, Stefani, Saunshi, Nikunj, Miryoosefi, Sobhan, Reddi, Sashank J., Kumar, Sanjiv
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911903680299008
author Karp, Stefani
Saunshi, Nikunj
Miryoosefi, Sobhan
Reddi, Sashank J.
Kumar, Sanjiv
author_facet Karp, Stefani
Saunshi, Nikunj
Miryoosefi, Sobhan
Reddi, Sashank J.
Kumar, Sanjiv
contents Recently, there has been increasing interest in efficient pretraining paradigms for training Transformer-based models. Several recent approaches use smaller models to initialize larger models in order to save computation (e.g., stacking and fusion). In this work, we study the fundamental question of how to select the best growing strategy from a given pool of growing strategies. Prior works have extensively focused on loss- and/or function-preserving behavior at initialization or simply performance at the end of training. Instead, we identify that behavior at initialization can be misleading as a predictor of final performance and present an alternative perspective based on early training dynamics, which we call "landscape-aware growing (LAG)". We perform extensive analysis of correlation of the final performance with performance in the initial steps of training and find early and more accurate predictions of the optimal growing strategy (i.e., with only a small "lag" after initialization). This perspective also motivates an adaptive strategy for gradual stacking.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Landscape-Aware Growing: The Power of a Little LAG
Karp, Stefani
Saunshi, Nikunj
Miryoosefi, Sobhan
Reddi, Sashank J.
Kumar, Sanjiv
Machine Learning
Computation and Language
Recently, there has been increasing interest in efficient pretraining paradigms for training Transformer-based models. Several recent approaches use smaller models to initialize larger models in order to save computation (e.g., stacking and fusion). In this work, we study the fundamental question of how to select the best growing strategy from a given pool of growing strategies. Prior works have extensively focused on loss- and/or function-preserving behavior at initialization or simply performance at the end of training. Instead, we identify that behavior at initialization can be misleading as a predictor of final performance and present an alternative perspective based on early training dynamics, which we call "landscape-aware growing (LAG)". We perform extensive analysis of correlation of the final performance with performance in the initial steps of training and find early and more accurate predictions of the optimal growing strategy (i.e., with only a small "lag" after initialization). This perspective also motivates an adaptive strategy for gradual stacking.
title Landscape-Aware Growing: The Power of a Little LAG
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2406.02469