Just One Layer Norm Guarantees Stable Extrapolation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ziomek, Juliusz, Whittle, George, Osborne, Michael A. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
par: Whittle, George, et autres
Publié: (2025)
par: Whittle, George, et autres
Publié: (2025)
Canonical Regularisation of Wide Feature-Learning Neural Networks
par: Whittle, George, et autres
Publié: (2026)
par: Whittle, George, et autres
Publié: (2026)
Time-Varying Gaussian Process Bandits with Unknown Prior
par: Ziomek, Juliusz, et autres
Publié: (2024)
par: Ziomek, Juliusz, et autres
Publié: (2024)
Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal
par: Ziomek, Juliusz, et autres
Publié: (2024)
par: Ziomek, Juliusz, et autres
Publié: (2024)
Open-Ended Task Discovery via Bayesian Optimization
par: Adachi, Masaki, et autres
Publié: (2026)
par: Adachi, Masaki, et autres
Publié: (2026)
Mean-Field Bayesian Optimisation
par: Steinberg, Petar, et autres
Publié: (2025)
par: Steinberg, Petar, et autres
Publié: (2025)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
par: Ziomek, Juliusz, et autres
Publié: (2026)
par: Ziomek, Juliusz, et autres
Publié: (2026)
Extrapolation Guarantees for Perturbation Modeling Under the Additive Latent Shift Assumption
par: von Kügelgen, Julius, et autres
Publié: (2025)
par: von Kügelgen, Julius, et autres
Publié: (2025)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
par: Chen, Chen, et autres
Publié: (2026)
par: Chen, Chen, et autres
Publié: (2026)
Geometry and Dynamics of LayerNorm
par: Riechers, Paul M.
Publié: (2024)
par: Riechers, Paul M.
Publié: (2024)
Stable GFlowNets with Probabilistic Guarantees
par: Lei, Zengxiang, et autres
Publié: (2026)
par: Lei, Zengxiang, et autres
Publié: (2026)
Informal Safety Guarantees for Simulated Optimizers Through Extrapolation from Partial Simulations
par: Marks, Luke
Publié: (2023)
par: Marks, Luke
Publié: (2023)
On the Role of Attention Masks and LayerNorm in Transformers
par: Wu, Xinyi, et autres
Publié: (2024)
par: Wu, Xinyi, et autres
Publié: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
par: Baroni, Luca, et autres
Publié: (2025)
par: Baroni, Luca, et autres
Publié: (2025)
TI-DeepONet: Learnable Time Integration for Stable Long-Term Extrapolation
par: Nayak, Dibyajyoti, et autres
Publié: (2025)
par: Nayak, Dibyajyoti, et autres
Publié: (2025)
Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
par: Grishina, Ekaterina, et autres
Publié: (2024)
par: Grishina, Ekaterina, et autres
Publié: (2024)
Optimal Flow Matching: Learning Straight Trajectories in Just One Step
par: Kornilov, Nikita, et autres
Publié: (2024)
par: Kornilov, Nikita, et autres
Publié: (2024)
Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation
par: Ousherovitch, Adam, et autres
Publié: (2026)
par: Ousherovitch, Adam, et autres
Publié: (2026)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
par: Wang, Puyu, et autres
Publié: (2023)
par: Wang, Puyu, et autres
Publié: (2023)
LayerNorm Induces Recency Bias in Transformer Decoders
par: Kim, Junu, et autres
Publié: (2025)
par: Kim, Junu, et autres
Publié: (2025)
Purification Of Contaminated Convolutional Neural Networks Via Robust Recovery: An Approach with Theoretical Guarantee in One-Hidden-Layer Case
par: Lu, Hanxiao, et autres
Publié: (2024)
par: Lu, Hanxiao, et autres
Publié: (2024)
SLaNC: Static LayerNorm Calibration
par: Salmani, Mahsa, et autres
Publié: (2024)
par: Salmani, Mahsa, et autres
Publié: (2024)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
par: Nichani, Eshaan, et autres
Publié: (2023)
par: Nichani, Eshaan, et autres
Publié: (2023)
Survival Kernets: Scalable and Interpretable Deep Kernel Survival Analysis with an Accuracy Guarantee
par: Chen, George H.
Publié: (2022)
par: Chen, George H.
Publié: (2022)
An Efficient Variant of One-Class SVM with Lifelong Online Learning Guarantees
par: Suk, Joe, et autres
Publié: (2025)
par: Suk, Joe, et autres
Publié: (2025)
On Logical Extrapolation for Mazes with Recurrent and Implicit Networks
par: Knutson, Brandon, et autres
Publié: (2024)
par: Knutson, Brandon, et autres
Publié: (2024)
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
par: Hsu, Hsin-Ling, et autres
Publié: (2026)
par: Hsu, Hsin-Ling, et autres
Publié: (2026)
Layer-wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep Learning
par: Lee, Sunwoo
Publié: (2025)
par: Lee, Sunwoo
Publié: (2025)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
par: Qian, Wenhao, et autres
Publié: (2025)
par: Qian, Wenhao, et autres
Publié: (2025)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
par: Delattre, Blaise, et autres
Publié: (2024)
par: Delattre, Blaise, et autres
Publié: (2024)
You can remove GPT2's LayerNorm by fine-tuning
par: Heimersheim, Stefan
Publié: (2024)
par: Heimersheim, Stefan
Publié: (2024)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
par: Dang, Quy-Anh, et autres
Publié: (2026)
par: Dang, Quy-Anh, et autres
Publié: (2026)
LayerNorm: A key component in parameter-efficient fine-tuning
par: ValizadehAslani, Taha, et autres
Publié: (2024)
par: ValizadehAslani, Taha, et autres
Publié: (2024)
Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence Guarantees
par: Nguyen, Thien Hang, et autres
Publié: (2024)
par: Nguyen, Thien Hang, et autres
Publié: (2024)
SEDGE: Structural Extrapolated Data Generation
par: Zhang, Kun, et autres
Publié: (2026)
par: Zhang, Kun, et autres
Publié: (2026)
On Model Extrapolation in Marginal Shapley Values
par: Rozenfeld, Ilya
Publié: (2024)
par: Rozenfeld, Ilya
Publié: (2024)
One-Stage Top-$k$ Learning-to-Defer: Score-Based Surrogates with Theoretical Guarantees
par: Montreuil, Yannis, et autres
Publié: (2025)
par: Montreuil, Yannis, et autres
Publié: (2025)
Mesa-Extrapolation: A Weave Position Encoding Method for Enhanced Extrapolation in LLMs
par: Ma, Xin, et autres
Publié: (2024)
par: Ma, Xin, et autres
Publié: (2024)
Neural Ordinary Differential Equations for Learning and Extrapolating System Dynamics Across Bifurcations
par: van Tegelen, Eva, et autres
Publié: (2025)
par: van Tegelen, Eva, et autres
Publié: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
par: He, Zhenyu, et autres
Publié: (2024)
par: He, Zhenyu, et autres
Publié: (2024)
Documents similaires
-
Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
par: Whittle, George, et autres
Publié: (2025) -
Canonical Regularisation of Wide Feature-Learning Neural Networks
par: Whittle, George, et autres
Publié: (2026) -
Time-Varying Gaussian Process Bandits with Unknown Prior
par: Ziomek, Juliusz, et autres
Publié: (2024) -
Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal
par: Ziomek, Juliusz, et autres
Publié: (2024) -
Open-Ended Task Discovery via Bayesian Optimization
par: Adachi, Masaki, et autres
Publié: (2026)