Geometry and Dynamics of LayerNorm
Fuente:
arXiv
Saved in:
| Main Author: | Riechers, Paul M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025)
by: Baroni, Luca, et al.
Published: (2025)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
LayerNorm: A key component in parameter-efficient fine-tuning
by: ValizadehAslani, Taha, et al.
Published: (2024)
by: ValizadehAslani, Taha, et al.
Published: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
by: Qian, Wenhao, et al.
Published: (2025)
by: Qian, Wenhao, et al.
Published: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm
by: Fan, Xiao, et al.
Published: (2025)
by: Fan, Xiao, et al.
Published: (2025)
Rank-1 LoRAs Encode Interpretable Reasoning Signals
by: Ward, Jake, et al.
Published: (2025)
by: Ward, Jake, et al.
Published: (2025)
Constrained belief updates explain geometric structures in transformer representations
by: Piotrowski, Mateusz, et al.
Published: (2025)
by: Piotrowski, Mateusz, et al.
Published: (2025)
Neural Collapse Dynamics: Depth, Activation, Regularisation, and Feature Norm Threshold
by: Rupa, Anamika Paul
Published: (2026)
by: Rupa, Anamika Paul
Published: (2026)
Neural networks leverage nominally quantum and post-quantum representations
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
Just One Layer Norm Guarantees Stable Extrapolation
by: Ziomek, Juliusz, et al.
Published: (2025)
by: Ziomek, Juliusz, et al.
Published: (2025)
The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
by: Musat, Tiberiu
Published: (2025)
by: Musat, Tiberiu
Published: (2025)
Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
by: Grishina, Ekaterina, et al.
Published: (2024)
by: Grishina, Ekaterina, et al.
Published: (2024)
Next-token pretraining implies in-context learning
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
Layer-wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep Learning
by: Lee, Sunwoo
Published: (2025)
by: Lee, Sunwoo
Published: (2025)
Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
by: Dang, Quy-Anh, et al.
Published: (2026)
by: Dang, Quy-Anh, et al.
Published: (2026)
MimicNorm: Weight Mean and Last BN Layer Mimic the Dynamic of Batch Normalization
by: Fei, Wen, et al.
Published: (2020)
by: Fei, Wen, et al.
Published: (2020)
Transformers represent belief state geometry in their residual stream
by: Shai, Adam S., et al.
Published: (2024)
by: Shai, Adam S., et al.
Published: (2024)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
Greedy-Gnorm: A Gradient Matrix Norm-Based Alternative to Attention Entropy for Head Pruning
by: Guo, Yuxi, et al.
Published: (2026)
by: Guo, Yuxi, et al.
Published: (2026)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
From Layers to Networks: Comparing Neural Representations via Diffusion Geometry
by: Khandait, Atharva, et al.
Published: (2026)
by: Khandait, Atharva, et al.
Published: (2026)
Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
Improving LLM Final Representations with Inter-Layer Geometry
by: Ulanovski, Tom, et al.
Published: (2026)
by: Ulanovski, Tom, et al.
Published: (2026)
Scalable Optimization in the Modular Norm
by: Large, Tim, et al.
Published: (2024)
by: Large, Tim, et al.
Published: (2024)
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
by: Cao, Tianxiao, et al.
Published: (2025)
by: Cao, Tianxiao, et al.
Published: (2025)
Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits
by: Suder, Piotr M., et al.
Published: (2025)
by: Suder, Piotr M., et al.
Published: (2025)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
by: Sim, Woo Seob, et al.
Published: (2026)
by: Sim, Woo Seob, et al.
Published: (2026)
FlashNorm: Fast Normalization for Transformers
by: Graef, Nils, et al.
Published: (2024)
by: Graef, Nils, et al.
Published: (2024)
Nuclear Norm Regularization for Deep Learning
by: Scarvelis, Christopher, et al.
Published: (2024)
by: Scarvelis, Christopher, et al.
Published: (2024)
Norm-Bounded Low-Rank Adaptation
by: Wang, Ruigang, et al.
Published: (2025)
by: Wang, Ruigang, et al.
Published: (2025)
Operationalising Rawlsian Ethics for Fairness in Norm-Learning Agents
by: Woodgate, Jessica, et al.
Published: (2024)
by: Woodgate, Jessica, et al.
Published: (2024)
Dynamic Switch Layers For Unsupervised Learning
by: Li, Haiguang, et al.
Published: (2024)
by: Li, Haiguang, et al.
Published: (2024)
Similar Items
-
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024) -
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
by: Baroni, Luca, et al.
Published: (2025) -
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024) -
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025) -
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
by: Chen, Chen, et al.
Published: (2026)