Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Fuente:
arXiv
Guardado en:
| Autor principal: | Verma, Lucky |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decoupled Weight Decay for Any $p$ Norm
por: Outmezguine, Nadav Joseph, et al.
Publicado: (2024)
por: Outmezguine, Nadav Joseph, et al.
Publicado: (2024)
Set-based Neural Network Encoding Without Weight Tying
por: Andreis, Bruno, et al.
Publicado: (2023)
por: Andreis, Bruno, et al.
Publicado: (2023)
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
por: Mukherji, Rishav, et al.
Publicado: (2024)
por: Mukherji, Rishav, et al.
Publicado: (2024)
Diagnostic Method for Hydropower Plant Condition-based Maintenance combining Autoencoder with Clustering Algorithms
por: Jad, Samy, et al.
Publicado: (2025)
por: Jad, Samy, et al.
Publicado: (2025)
Evolved Sample Weights for Bias Mitigation: Effectiveness Depends on the Fairness Objective
por: Saini, Anil K., et al.
Publicado: (2025)
por: Saini, Anil K., et al.
Publicado: (2025)
Can We Optimize Deep RL Policy Weights as Trajectory Modeling?
por: Tang, Hongyao
Publicado: (2025)
por: Tang, Hongyao
Publicado: (2025)
Recursive Dynamics in Fast-Weights Homeostatic Reentry Networks: Toward Reflective Intelligence
por: Chae, B. G.
Publicado: (2025)
por: Chae, B. G.
Publicado: (2025)
Attending to Graph Transformers
por: Müller, Luis, et al.
Publicado: (2023)
por: Müller, Luis, et al.
Publicado: (2023)
Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks
por: Xiao, Mingqing, et al.
Publicado: (2024)
por: Xiao, Mingqing, et al.
Publicado: (2024)
Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats
por: Heilper, Anat, et al.
Publicado: (2025)
por: Heilper, Anat, et al.
Publicado: (2025)
Structure Development in List-Sorting Transformers
por: Urdshals, Einar, et al.
Publicado: (2025)
por: Urdshals, Einar, et al.
Publicado: (2025)
Investigating Recurrent Transformers with Dynamic Halt
por: Chowdhury, Jishnu Ray, et al.
Publicado: (2024)
por: Chowdhury, Jishnu Ray, et al.
Publicado: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
Spiking Point Transformer for Point Cloud Classification
por: Wu, Peixi, et al.
Publicado: (2025)
por: Wu, Peixi, et al.
Publicado: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
por: Csordás, Róbert, et al.
Publicado: (2024)
por: Csordás, Róbert, et al.
Publicado: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
por: Kirsch, Louis, et al.
Publicado: (2022)
por: Kirsch, Louis, et al.
Publicado: (2022)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
por: Putra, Rachmad Vidya Wicaksana, et al.
Publicado: (2025)
por: Putra, Rachmad Vidya Wicaksana, et al.
Publicado: (2025)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
por: Zhang, Huizhe, et al.
Publicado: (2024)
por: Zhang, Huizhe, et al.
Publicado: (2024)
Interpretable Fine-Gray Deep Survival Model for Competing Risks: Predicting Post-Discharge Foot Complications for Diabetic Patients in Ontario
por: Ramachandram, Dhanesh, et al.
Publicado: (2025)
por: Ramachandram, Dhanesh, et al.
Publicado: (2025)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
por: Kosowski, Adrian, et al.
Publicado: (2025)
por: Kosowski, Adrian, et al.
Publicado: (2025)
Why "classic" Transformers are shallow and how to make them go deep
por: Yu, Yueyao, et al.
Publicado: (2023)
por: Yu, Yueyao, et al.
Publicado: (2023)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
por: Smithline, Gabriel, et al.
Publicado: (2026)
por: Smithline, Gabriel, et al.
Publicado: (2026)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
por: Meng, Fanfei, et al.
Publicado: (2023)
por: Meng, Fanfei, et al.
Publicado: (2023)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
por: Bergsma, Shane, et al.
Publicado: (2025)
por: Bergsma, Shane, et al.
Publicado: (2025)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
por: Chaudhary, Siddharth
Publicado: (2025)
por: Chaudhary, Siddharth
Publicado: (2025)
ActPC-Geom: Towards Scalable Online Neural-Symbolic Learning via Accelerating Active Predictive Coding with Information Geometry & Diverse Cognitive Mechanisms
por: Goertzel, Ben
Publicado: (2025)
por: Goertzel, Ben
Publicado: (2025)
Decoding Listeners Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
por: Lin, Zheyuan, et al.
Publicado: (2025)
por: Lin, Zheyuan, et al.
Publicado: (2025)
Learning to Forget: Continual Learning with Adaptive Weight Decay
por: Ramesh, Aditya A., et al.
Publicado: (2026)
por: Ramesh, Aditya A., et al.
Publicado: (2026)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
por: Musto, Henry, et al.
Publicado: (2024)
por: Musto, Henry, et al.
Publicado: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2024)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2024)
SNN-Based Online Learning of Concepts and Action Laws in an Open World
por: Grimaud, Christel, et al.
Publicado: (2024)
por: Grimaud, Christel, et al.
Publicado: (2024)
Comply: Learning Sentences with Complex Weights inspired by Fruit Fly Olfaction
por: Figueroa, Alexei, et al.
Publicado: (2025)
por: Figueroa, Alexei, et al.
Publicado: (2025)
Datum-wise Transformer for Synthetic Tabular Data Detection in the Wild
por: Kindji, G. Charbel N., et al.
Publicado: (2025)
por: Kindji, G. Charbel N., et al.
Publicado: (2025)
GLU Attention Improve Transformer
por: Wang, Zehao
Publicado: (2025)
por: Wang, Zehao
Publicado: (2025)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
por: Smith, Ethan
Publicado: (2026)
por: Smith, Ethan
Publicado: (2026)
A Transformer-based Neural Architecture Search Method
por: Wang, Shang, et al.
Publicado: (2025)
por: Wang, Shang, et al.
Publicado: (2025)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
por: Lee, Philip Heejun
Publicado: (2025)
por: Lee, Philip Heejun
Publicado: (2025)
Learning under noisy supervision is governed by a feedback-truth gap
por: Schonfeld, Elan, et al.
Publicado: (2026)
por: Schonfeld, Elan, et al.
Publicado: (2026)
Ejemplares similares
-
Decoupled Weight Decay for Any $p$ Norm
por: Outmezguine, Nadav Joseph, et al.
Publicado: (2024) -
Set-based Neural Network Encoding Without Weight Tying
por: Andreis, Bruno, et al.
Publicado: (2023) -
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
por: Mukherji, Rishav, et al.
Publicado: (2024) -
Diagnostic Method for Hydropower Plant Condition-based Maintenance combining Autoencoder with Clustering Algorithms
por: Jad, Samy, et al.
Publicado: (2025) -
Evolved Sample Weights for Bias Mitigation: Effectiveness Depends on the Fairness Objective
por: Saini, Anil K., et al.
Publicado: (2025)