GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Tianhao, Xu, Xin, Liu, Zijing, Li, Pengxiang, Song, Xinyuan, Jaiswal, Ajay Kumar, Zhang, Fan, Hu, Jishan, Wang, Yang, Chen, Hao, Diao, Shizhe, Liu, Shiwei, Li, Yu, Yin, Lu, Yang, Can
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916824496472064
author Chen, Tianhao
Xu, Xin
Liu, Zijing
Li, Pengxiang
Song, Xinyuan
Jaiswal, Ajay Kumar
Zhang, Fan
Hu, Jishan
Wang, Yang
Chen, Hao
Diao, Shizhe
Liu, Shiwei
Li, Yu
Yin, Lu
Yang, Can
author_facet Chen, Tianhao
Xu, Xin
Liu, Zijing
Li, Pengxiang
Song, Xinyuan
Jaiswal, Ajay Kumar
Zhang, Fan
Hu, Jishan
Wang, Yang
Chen, Hao
Diao, Shizhe
Liu, Shiwei
Li, Yu
Yin, Lu
Yang, Can
contents Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
Chen, Tianhao
Xu, Xin
Liu, Zijing
Li, Pengxiang
Song, Xinyuan
Jaiswal, Ajay Kumar
Zhang, Fan
Hu, Jishan
Wang, Yang
Chen, Hao
Diao, Shizhe
Liu, Shiwei
Li, Yu
Yin, Lu
Yang, Can
Machine Learning
Computation and Language
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
title GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.22049