HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhuo, Zhijian, Zeng, Yutao, Wang, Ya, Zhang, Sijun, Yang, Jian, Li, Xiaoqing, Zhou, Xun, Ma, Jinwen
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917132585926656
author Zhuo, Zhijian
Zeng, Yutao
Wang, Ya
Zhang, Sijun
Yang, Jian
Li, Xiaoqing
Zhou, Xun
Ma, Jinwen
author_facet Zhuo, Zhijian
Zeng, Yutao
Wang, Ya
Zhang, Sijun
Yang, Jian
Li, Xiaoqing
Zhou, Xun
Ma, Jinwen
contents Transformers have become the de facto architecture for a wide range of machine learning tasks, particularly in large language models (LLMs). Despite their remarkable performance, many challenges remain in training deep transformer networks, especially regarding the position of the layer normalization. While Pre-Norm structures facilitate more stable training owing to their stronger identity path, they often lead to suboptimal performance compared to Post-Norm. In this paper, we propose $\textbf{HybridNorm}$, a simple yet effective hybrid normalization strategy that integrates the advantages of both Pre-Norm and Post-Norm. Specifically, HybridNorm employs QKV normalization within the attention mechanism and Post-Norm in the feed-forward network (FFN) of each transformer block. We provide both theoretical insights and empirical evidence to demonstrate that HybridNorm improves the gradient flow and the model robustness. Extensive experiments on large-scale transformer models, including both dense and sparse variants, show that HybridNorm consistently outperforms both Pre-Norm and Post-Norm approaches across multiple benchmarks. These findings highlight the potential of HybridNorm as a more stable and effective technique for improving the training and performance of deep transformer models. Code is available at https://github.com/BryceZhuo/HybridNorm.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
Zhuo, Zhijian
Zeng, Yutao
Wang, Ya
Zhang, Sijun
Yang, Jian
Li, Xiaoqing
Zhou, Xun
Ma, Jinwen
Computation and Language
Artificial Intelligence
Machine Learning
Transformers have become the de facto architecture for a wide range of machine learning tasks, particularly in large language models (LLMs). Despite their remarkable performance, many challenges remain in training deep transformer networks, especially regarding the position of the layer normalization. While Pre-Norm structures facilitate more stable training owing to their stronger identity path, they often lead to suboptimal performance compared to Post-Norm. In this paper, we propose $\textbf{HybridNorm}$, a simple yet effective hybrid normalization strategy that integrates the advantages of both Pre-Norm and Post-Norm. Specifically, HybridNorm employs QKV normalization within the attention mechanism and Post-Norm in the feed-forward network (FFN) of each transformer block. We provide both theoretical insights and empirical evidence to demonstrate that HybridNorm improves the gradient flow and the model robustness. Extensive experiments on large-scale transformer models, including both dense and sparse variants, show that HybridNorm consistently outperforms both Pre-Norm and Post-Norm approaches across multiple benchmarks. These findings highlight the potential of HybridNorm as a more stable and effective technique for improving the training and performance of deep transformer models. Code is available at https://github.com/BryceZhuo/HybridNorm.
title HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.04598