Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Chao, Li, Bei, Zhang, Jiaqi, Liu, Xinyu, Fan, Yuchun, Lyu, Linkun, Chen, Xin, Wang, Jingang, Xiao, Tong, Pei, Peng, Cai, Xunliang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2601.22580
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918314735828992
author Wang, Chao
Li, Bei
Zhang, Jiaqi
Liu, Xinyu
Fan, Yuchun
Lyu, Linkun
Chen, Xin
Wang, Jingang
Xiao, Tong
Pei, Peng
Cai, Xunliang
author_facet Wang, Chao
Li, Bei
Zhang, Jiaqi
Liu, Xinyu
Fan, Yuchun
Lyu, Linkun
Chen, Xin
Wang, Jingang
Xiao, Tong
Pei, Peng
Cai, Xunliang
contents The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ``PreNorm'' architecture ensures training stability at the cost of potential performance degradation in deep models, while the ``PostNorm'' architecture offers strong performance but suffers from severe training instability. In this work, we propose SpanNorm, a novel technique designed to resolve this dilemma by integrating the strengths of both paradigms. Structurally, SpanNorm establishes a clean residual connection that spans the entire transformer block to stabilize signal propagation, while employing a PostNorm-style computation that normalizes the aggregated output to enhance model performance. We provide a theoretical analysis demonstrating that SpanNorm, combined with a principled scaling strategy, maintains bounded signal variance throughout the network, preventing the gradient issues that plague PostNorm models, and also alleviating the representation collapse of PreNorm. Empirically, SpanNorm consistently outperforms standard normalization schemes in both dense and Mixture-of-Experts (MoE) scenarios, paving the way for more powerful and stable Transformer architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22580
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
Wang, Chao
Li, Bei
Zhang, Jiaqi
Liu, Xinyu
Fan, Yuchun
Lyu, Linkun
Chen, Xin
Wang, Jingang
Xiao, Tong
Pei, Peng
Cai, Xunliang
Computation and Language
Artificial Intelligence
Machine Learning
The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ``PreNorm'' architecture ensures training stability at the cost of potential performance degradation in deep models, while the ``PostNorm'' architecture offers strong performance but suffers from severe training instability. In this work, we propose SpanNorm, a novel technique designed to resolve this dilemma by integrating the strengths of both paradigms. Structurally, SpanNorm establishes a clean residual connection that spans the entire transformer block to stabilize signal propagation, while employing a PostNorm-style computation that normalizes the aggregated output to enhance model performance. We provide a theoretical analysis demonstrating that SpanNorm, combined with a principled scaling strategy, maintains bounded signal variance throughout the network, preventing the gradient issues that plague PostNorm models, and also alleviating the representation collapse of PreNorm. Empirically, SpanNorm consistently outperforms standard normalization schemes in both dense and Mixture-of-Experts (MoE) scenarios, paving the way for more powerful and stable Transformer architectures.
title SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.22580