Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Gray, Gavia, Tiwari, Aman, Bergsma, Shane, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Identifying Policy Gradient Subspaces
by: Schneider, Jan, et al.
Published: (2024)
by: Schneider, Jan, et al.
Published: (2024)
Low-Rank Adapters Initialization via Gradient Surgery for Continual Learning
by: Pasquali, Joana, et al.
Published: (2026)
by: Pasquali, Joana, et al.
Published: (2026)
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
by: Pawar, Urvi, et al.
Published: (2025)
by: Pawar, Urvi, et al.
Published: (2025)
Difference Rewards Policy Gradients
by: Castellini, Jacopo, et al.
Published: (2020)
by: Castellini, Jacopo, et al.
Published: (2020)
Predictable Gradient Manifolds in Deep Learning: Temporal Path-Length and Intrinsic Rank as a Complexity Regime
by: Calvo, Anherutowa
Published: (2026)
by: Calvo, Anherutowa
Published: (2026)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
by: Katakam, Raj Kiran Gupta
Published: (2026)
by: Katakam, Raj Kiran Gupta
Published: (2026)
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
NoRIN: Backbone-Adaptive Reversible Normalization for Time-Series Forecasting
by: Zhang, Shun, et al.
Published: (2026)
by: Zhang, Shun, et al.
Published: (2026)
Composing Linear Layers from Irreducibles
by: Pence, Travis, et al.
Published: (2025)
by: Pence, Travis, et al.
Published: (2025)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
AGOP-IxG: A Gradient Covariance Filter for Local Feature Attribution on Tabular Data, with a Controlled Benchmark
by: Katakam, Raj Kiran Gupta
Published: (2026)
by: Katakam, Raj Kiran Gupta
Published: (2026)
Adaptive Epsilon Adversarial Training for Robust Gravitational Wave Parameter Estimation Using Normalizing Flows
by: Yang, Yiqian, et al.
Published: (2024)
by: Yang, Yiqian, et al.
Published: (2024)
Real-Time Pulsatile Flow Prediction for Realistic, Diverse Intracranial Aneurysm Morphologies using a Graph Transformer and Steady-Flow Data Augmentation
by: Sheng, Yiying, et al.
Published: (2026)
by: Sheng, Yiying, et al.
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Energy and Memory-Efficient Federated Learning With Ordered Layer Freezing
by: Niu, Ziru, et al.
Published: (2025)
by: Niu, Ziru, et al.
Published: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
by: Heyman, Alex, et al.
Published: (2025)
by: Heyman, Alex, et al.
Published: (2025)
Scaling Laws for Neural Material Models
by: Trikha, Akshay, et al.
Published: (2025)
by: Trikha, Akshay, et al.
Published: (2025)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Learning Transferable Predictability Representations
by: Goswami, Diyali, et al.
Published: (2026)
by: Goswami, Diyali, et al.
Published: (2026)
Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control
by: Paulus, Anselm, et al.
Published: (2025)
by: Paulus, Anselm, et al.
Published: (2025)
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Fine-grained Attention in Hierarchical Transformers for Tabular Time-series
by: Azorin, Raphael, et al.
Published: (2024)
by: Azorin, Raphael, et al.
Published: (2024)
FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models
by: Polly, Fabien
Published: (2026)
by: Polly, Fabien
Published: (2026)
MotherNet: Fast Training and Inference via Hyper-Network Transformers
by: Müller, Andreas, et al.
Published: (2023)
by: Müller, Andreas, et al.
Published: (2023)
Informed Priors for Knowledge Integration in Trajectory Prediction
by: Schlauch, Christian, et al.
Published: (2022)
by: Schlauch, Christian, et al.
Published: (2022)
Explaining Temporal Graph Predictions With Shapley Values
by: Sussek, Lea-Marie, et al.
Published: (2026)
by: Sussek, Lea-Marie, et al.
Published: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Faster Predictive Coding Networks via Better Initialization
by: Pinchetti, Luca, et al.
Published: (2026)
by: Pinchetti, Luca, et al.
Published: (2026)
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
Kolmogorov Arnold Networks and Multi-Layer Perceptrons: A Paradigm Shift in Neural Modelling
by: Gaonkar, Aradhya, et al.
Published: (2026)
by: Gaonkar, Aradhya, et al.
Published: (2026)
Pre-Ictal Seizure Prediction Using Personalized Deep Learning
by: Jaddu, Shriya, et al.
Published: (2024)
by: Jaddu, Shriya, et al.
Published: (2024)
Measuring In-Context Computation Complexity via Hidden State Prediction
by: Herrmann, Vincent, et al.
Published: (2025)
by: Herrmann, Vincent, et al.
Published: (2025)
Variational Autoencoder with Normalizing flow for X-ray spectral fitting
by: Redmen, Fiona, et al.
Published: (2026)
by: Redmen, Fiona, et al.
Published: (2026)
Asynchronous Stochastic Gradient Descent with Decoupled Backpropagation and Layer-Wise Updates
by: Fokam, Cabrel Teguemne, et al.
Published: (2024)
by: Fokam, Cabrel Teguemne, et al.
Published: (2024)
Versatile Ordering Network: An Attention-based Neural Network for Ordering Across Scales and Quality Metrics
by: Yu, Zehua, et al.
Published: (2024)
by: Yu, Zehua, et al.
Published: (2024)
Similar Items
-
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
by: Steele, Brady
Published: (2026) -
Identifying Policy Gradient Subspaces
by: Schneider, Jan, et al.
Published: (2024) -
Low-Rank Adapters Initialization via Gradient Surgery for Continual Learning
by: Pasquali, Joana, et al.
Published: (2026) -
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
by: Pawar, Urvi, et al.
Published: (2025) -
Difference Rewards Policy Gradients
by: Castellini, Jacopo, et al.
Published: (2020)