Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Pengxiang, Yin, Lu, Liu, Shiwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Point-LN: A Lightweight Framework for Efficient Point Cloud Classification Using Non-Parametric Positional Encoding
by: Mohammadi, Marzieh, et al.
Published: (2025)
by: Mohammadi, Marzieh, et al.
Published: (2025)
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
The Curse of Depth in Large Language Models
by: Sun, Wenfang, et al.
Published: (2025)
by: Sun, Wenfang, et al.
Published: (2025)
DP-aware AdaLN-Zero: Taming Conditioning-Induced Heavy-Tailed Gradients in Differentially Private Diffusion
by: Huang, Tao, et al.
Published: (2026)
by: Huang, Tao, et al.
Published: (2026)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
by: Li, Jiaxi, et al.
Published: (2026)
by: Li, Jiaxi, et al.
Published: (2026)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)
by: Nepal, Aadim, et al.
Published: (2025)
MLCopilot: Unleashing the Power of Large Language Models in Solving Machine Learning Tasks
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
Power Plays: Unleashing Machine Learning Magic in Smart Grids
by: Rashid, Abdur, et al.
Published: (2024)
by: Rashid, Abdur, et al.
Published: (2024)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
by: Zhang, Gengwei, et al.
Published: (2024)
by: Zhang, Gengwei, et al.
Published: (2024)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026)
by: He, Di, et al.
Published: (2026)
Deeper Insights into Deep Graph Convolutional Networks: Stability and Generalization
by: Yang, Guangrui, et al.
Published: (2024)
by: Yang, Guangrui, et al.
Published: (2024)
ResNets Are Deeper Than You Think
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
OledFL: Unleashing the Potential of Decentralized Federated Learning via Opposite Lookahead Enhancement
by: Li, Qinglun, et al.
Published: (2024)
by: Li, Qinglun, et al.
Published: (2024)
CSformer: Combining Channel Independence and Mixing for Robust Multivariate Time Series Forecasting
by: Wang, Haoxin, et al.
Published: (2023)
by: Wang, Haoxin, et al.
Published: (2023)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Digging Deeper: Learning Multi-Level Concept Hierarchies
by: Hill, Oscar, et al.
Published: (2026)
by: Hill, Oscar, et al.
Published: (2026)
The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
by: Su, Zelal, et al.
Published: (2026)
by: Su, Zelal, et al.
Published: (2026)
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models
by: Lv, Yaojia, et al.
Published: (2024)
by: Lv, Yaojia, et al.
Published: (2024)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
by: Huang, Nai-Chieh, et al.
Published: (2023)
by: Huang, Nai-Chieh, et al.
Published: (2023)
Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning
by: Piccoli, Elia, et al.
Published: (2025)
by: Piccoli, Elia, et al.
Published: (2025)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
Reconstructing Deep Neural Networks: Unleashing the Optimization Potential of Natural Gradient Descent
by: Liu, Weihua, et al.
Published: (2024)
by: Liu, Weihua, et al.
Published: (2024)
Self-Supervised Pre-Training for Precipitation Post-Processor
by: An, Sojung, et al.
Published: (2023)
by: An, Sojung, et al.
Published: (2023)
Sparser, Better, Deeper, Stronger: Improving Sparse Training with Exact Orthogonal Initialization
by: Nowak, Aleksandra Irena, et al.
Published: (2024)
by: Nowak, Aleksandra Irena, et al.
Published: (2024)
EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
LN-Gen: Rectal Lymph Nodes Generation via Anatomical Features
by: Guo, Weidong, et al.
Published: (2024)
by: Guo, Weidong, et al.
Published: (2024)
Lecture Notes in Loop Quantum Gravity. LN4: Hamiltonian framework
by: Fatibene, Lorenzo, et al.
Published: (2026)
by: Fatibene, Lorenzo, et al.
Published: (2026)
FinLN: 3D Fin Lithium Niobate Acoustic Resonators
by: Ni, Haorui, et al.
Published: (2026)
by: Ni, Haorui, et al.
Published: (2026)
A CONCEPÇÃO DE SUJEITO COMO (LN)VIABILIZADORA DA APRENDIZAGEM
by: Marianne Montenegro Stolzmann Mendes Ribeiro
Published: (2007)
by: Marianne Montenegro Stolzmann Mendes Ribeiro
Published: (2007)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
Graph Attention Networks Unleashed: A Fast and Explainable Vulnerability Assessment Framework for Microgrids
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Similar Items
-
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025) -
Point-LN: A Lightweight Framework for Efficient Point Cloud Classification Using Non-Parametric Positional Encoding
by: Mohammadi, Marzieh, et al.
Published: (2025) -
Outlier-weighed Layerwise Sampling for LLM Fine-tuning
by: Li, Pengxiang, et al.
Published: (2024) -
Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series
by: Zhang, Weijia, et al.
Published: (2024) -
The Curse of Depth in Large Language Models
by: Sun, Wenfang, et al.
Published: (2025)