Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Chen, Wei, Lai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
by: Zhuo, Zhijian, et al.
Published: (2025)
by: Zhuo, Zhijian, et al.
Published: (2025)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
by: Lai, Yuhang, et al.
Published: (2026)
by: Lai, Yuhang, et al.
Published: (2026)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
by: Ma, Xin, et al.
Published: (2026)
by: Ma, Xin, et al.
Published: (2026)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
by: London, Charles, et al.
Published: (2025)
by: London, Charles, et al.
Published: (2025)
Interactive and Expressive Code-Augmented Planning with Large Language Models
by: Liu, Anthony Z., et al.
Published: (2024)
by: Liu, Anthony Z., et al.
Published: (2024)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
An Algebraic View of the Expressivity of Recurrent Language Models
by: Nowak, Franz, et al.
Published: (2026)
by: Nowak, Franz, et al.
Published: (2026)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
by: Chen, Yan-Lun, et al.
Published: (2025)
by: Chen, Yan-Lun, et al.
Published: (2025)
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
The Expressive Power of Low-Rank Adaptation
by: Zeng, Yuchen, et al.
Published: (2023)
by: Zeng, Yuchen, et al.
Published: (2023)
Decoding Linguistic Nuances in Mental Health Text Classification Using Expressive Narrative Stories
by: Tang, Jinwen, et al.
Published: (2024)
by: Tang, Jinwen, et al.
Published: (2024)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
by: Sun, Libo, et al.
Published: (2026)
by: Sun, Libo, et al.
Published: (2026)
Olmo Hybrid: From Theory to Practice and Back
by: Merrill, William, et al.
Published: (2026)
by: Merrill, William, et al.
Published: (2026)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026)
by: Brösamle, Moritz, et al.
Published: (2026)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
by: Wong, Michel, et al.
Published: (2025)
by: Wong, Michel, et al.
Published: (2025)
Efficient Post-Training Pruning of Large Language Models with Statistical Correction
by: Yu, Peiqi, et al.
Published: (2026)
by: Yu, Peiqi, et al.
Published: (2026)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
by: Sun, Linzhuang, et al.
Published: (2024)
by: Sun, Linzhuang, et al.
Published: (2024)
FBQuant: FeedBack Quantization for Large Language Models
by: Liu, Yijiang, et al.
Published: (2025)
by: Liu, Yijiang, et al.
Published: (2025)
Similar Items
-
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025) -
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024) -
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024) -
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024) -
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)