Parallel Layer Normalization for Universal Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Yunhao, Guo, Yuxin, Liu, Yuhe, Sun, Wenxin, Luo, Jie, Wu, Wenjun, Huang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm
by: Guo, Yuxin, et al.
Published: (2026)
by: Guo, Yuxin, et al.
Published: (2026)
On the Nonlinearity of Layer Normalization
by: Ni, Yunhao, et al.
Published: (2024)
by: Ni, Yunhao, et al.
Published: (2024)
RLLaVA: An RL-central Framework for Language and Vision Assistants
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
Transformers Meet In-Context Learning: A Universal Approximation Theory
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Universal Approximation Theorem for a Single-Layer Transformer
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Looped Transformers with Layer Normalization Provably Learn the Power Method
by: Wu, Lyumin, et al.
Published: (2026)
by: Wu, Lyumin, et al.
Published: (2026)
Approximate Top-$k$ for Increased Parallelism
by: Key, Oscar, et al.
Published: (2024)
by: Key, Oscar, et al.
Published: (2024)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Constructive Universal Approximation and Sure Convergence for Multi-Layer Neural Networks
by: Chi, Chien-Ming
Published: (2025)
by: Chi, Chien-Ming
Published: (2025)
Cauchy-Schwarz Fairness Regularizer
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
Learning in Compact Spaces with Approximately Normalized Transformer
by: Franke, Jörg K. H., et al.
Published: (2025)
by: Franke, Jörg K. H., et al.
Published: (2025)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Unbiased Online Curvature Approximation for Regularized Graph Continual Learning
by: Yin, Jie, et al.
Published: (2025)
by: Yin, Jie, et al.
Published: (2025)
From Universal Approximation Theorem to Tropical Geometry of Multi-Layer Perceptrons
by: Chu, Yi-Shan, et al.
Published: (2025)
by: Chu, Yi-Shan, et al.
Published: (2025)
Modulate Your Spectrum in Self-Supervised Learning
by: Weng, Xi, et al.
Published: (2023)
by: Weng, Xi, et al.
Published: (2023)
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Decomposition-based multi-scale transformer framework for time series anomaly detection
by: Zhang, Wenxin, et al.
Published: (2025)
by: Zhang, Wenxin, et al.
Published: (2025)
Neural Approximation and Its Applications
by: Wu, Wei-Hao, et al.
Published: (2026)
by: Wu, Wei-Hao, et al.
Published: (2026)
TimeGMM: Single-Pass Probabilistic Forecasting via Adaptive Gaussian Mixture Models with Reversible Normalization
by: Liu, Lei, et al.
Published: (2026)
by: Liu, Lei, et al.
Published: (2026)
Enabling Group Fairness in Graph Unlearning via Bi-level Debiasing
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
A Semantic Partitioning Method for Large-Scale Training of Knowledge Graph Embeddings
by: Bai, Yuhe
Published: (2025)
by: Bai, Yuhe
Published: (2025)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming
by: Lin, Hao, et al.
Published: (2023)
by: Lin, Hao, et al.
Published: (2023)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
by: Choi, Dawon, et al.
Published: (2026)
by: Choi, Dawon, et al.
Published: (2026)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
by: Wu, Yize, et al.
Published: (2025)
by: Wu, Yize, et al.
Published: (2025)
Are Linear Regression Models White Box and Interpretable?
by: Salih, Ahmed M, et al.
Published: (2024)
by: Salih, Ahmed M, et al.
Published: (2024)
The Final Layer Holds the Key: A Unified and Efficient GNN Calibration Framework
by: Huang, Jincheng, et al.
Published: (2025)
by: Huang, Jincheng, et al.
Published: (2025)
Reservoir Static Property Estimation Using Nearest-Neighbor Neural Network
by: Wang, Yuhe
Published: (2024)
by: Wang, Yuhe
Published: (2024)
PEVLM: Parallel Encoding for Vision-Language Models
by: Kang, Letian, et al.
Published: (2025)
by: Kang, Letian, et al.
Published: (2025)
Dynamic Action Interpolation: A Universal Approach for Accelerating Reinforcement Learning with Expert Guidance
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Understanding the Role of Layer Normalization in Label-Skewed Federated Learning
by: Zhang, Guojun, et al.
Published: (2023)
by: Zhang, Guojun, et al.
Published: (2023)
SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer
by: Takida, Yuhta, et al.
Published: (2023)
by: Takida, Yuhta, et al.
Published: (2023)
SNLP: Layer-Parallel Inference via Structured Newton Corrections
by: Han, Ligong, et al.
Published: (2026)
by: Han, Ligong, et al.
Published: (2026)
Gradient Flow Drifting: Generative Modeling via Wasserstein Gradient Flows of KDE-Approximated Divergences
by: Cao, Jiarui, et al.
Published: (2026)
by: Cao, Jiarui, et al.
Published: (2026)
Similar Items
-
Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm
by: Guo, Yuxin, et al.
Published: (2026) -
On the Nonlinearity of Layer Normalization
by: Ni, Yunhao, et al.
Published: (2024) -
RLLaVA: An RL-central Framework for Language and Vision Assistants
by: Zhao, Lei, et al.
Published: (2025) -
An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
by: Wang, Li, et al.
Published: (2025) -
Transformers Meet In-Context Learning: A Universal Approximation Theory
by: Li, Gen, et al.
Published: (2025)