Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Baroni, Luca, Khara, Galvin, Schaeffer, Joachim, Subkhankulov, Marat, Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024)
by: Riechers, Paul M.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Optimal Scaling Needs Optimal Norm
by: Filatov, Oleg, et al.
Published: (2025)
by: Filatov, Oleg, et al.
Published: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Unpacking the Layers: Exploring Self-Disclosure Norms, Engagement Dynamics, and Privacy Implications
by: Haq, Ehsan-Ul, et al.
Published: (2025)
by: Haq, Ehsan-Ul, et al.
Published: (2025)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
by: Hatua, Amartya
Published: (2025)
by: Hatua, Amartya
Published: (2025)
"Don't Look, But I Know You Do": Norms and Observer Effects in Shared LLM Accounts
by: Song, Ji Eun, et al.
Published: (2026)
by: Song, Ji Eun, et al.
Published: (2026)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
Just One Layer Norm Guarantees Stable Extrapolation
by: Ziomek, Juliusz, et al.
Published: (2025)
by: Ziomek, Juliusz, et al.
Published: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025)
by: Braun, Dan, et al.
Published: (2025)
Don't Interpret ‘No’ As ‘Never’
Published: (2024)
Published: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
Evolution of SAE Features Across Layers in LLMs
by: Balcells, Daniel, et al.
Published: (2024)
by: Balcells, Daniel, et al.
Published: (2024)
Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
by: Grishina, Ekaterina, et al.
Published: (2024)
by: Grishina, Ekaterina, et al.
Published: (2024)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
by: Bushnaq, Lucius, et al.
Published: (2024)
by: Bushnaq, Lucius, et al.
Published: (2024)
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
by: Wang, Chenlong, et al.
Published: (2025)
by: Wang, Chenlong, et al.
Published: (2025)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
by: Dang, Quy-Anh, et al.
Published: (2026)
by: Dang, Quy-Anh, et al.
Published: (2026)
Love Don't Need a Reason
by: Jones, Matthew J.
Published: (2020)
by: Jones, Matthew J.
Published: (2020)
Teenagers Are Not Luggage: They Don't Need Handling.
by: Sullivan, Edward T.
Published: (2001)
by: Sullivan, Edward T.
Published: (2001)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
The Transformative Potential of Freelance Work: Positive Implications of a Global Norm
by: Kareemuddin, Fahad Kareemuddin
Published: (2025)
by: Kareemuddin, Fahad Kareemuddin
Published: (2025)
Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs
by: Shalankin, M.
Published: (2026)
by: Shalankin, M.
Published: (2026)
Who Says We Don't Need Catalogers?
by: Bishoff, Lizbeth J.
Published: (1987)
by: Bishoff, Lizbeth J.
Published: (1987)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
Chapter 4 Norms and Norm Contestation
by: Orchard, Phil, et al.
Published: (2022)
by: Orchard, Phil, et al.
Published: (2022)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)
by: Fei, Wen, et al.
Published: (2020)
Layer-wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep Learning
by: Lee, Sunwoo
Published: (2025)
by: Lee, Sunwoo
Published: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Similar Items
-
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024) -
Geometry and Dynamics of LayerNorm
by: Riechers, Paul M.
Published: (2024) -
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024) -
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025) -
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)