Saved in:
| Main Authors: | Noci, Lorenzo, Meterez, Alexandru, Hofmann, Thomas, Orvieto, Antonio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.17457 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
by: Belloni, Annalisa, et al.
Published: (2026)
by: Belloni, Annalisa, et al.
Published: (2026)
The Optimization Landscape of SGD Across the Feature Learning Strength
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Understanding and Minimising Outlier Features in Neural Network Training
by: He, Bobby, et al.
Published: (2024)
by: He, Bobby, et al.
Published: (2024)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
How Good is a Single Basin?
by: Lion, Kai, et al.
Published: (2024)
by: Lion, Kai, et al.
Published: (2024)
An Uncertainty Principle for Linear Recurrent Neural Networks
by: François, Alexandre, et al.
Published: (2025)
by: François, Alexandre, et al.
Published: (2025)
The Importance of Being Lazy: Scaling Limits of Continual Learning
by: Graldi, Jacopo, et al.
Published: (2025)
by: Graldi, Jacopo, et al.
Published: (2025)
Recurrent Distance Filtering for Graph Representation Learning
by: Ding, Yuhui, et al.
Published: (2023)
by: Ding, Yuhui, et al.
Published: (2023)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Adam Simplified: Bias Correction Debunked
by: Laing, Sam, et al.
Published: (2025)
by: Laing, Sam, et al.
Published: (2025)
Revisiting associative recall in modern recurrent models
by: Okpekpe, Destiny, et al.
Published: (2025)
by: Okpekpe, Destiny, et al.
Published: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
by: Dong, Yihe, et al.
Published: (2025)
by: Dong, Yihe, et al.
Published: (2025)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
In Search of Adam's Secret Sauce
by: Orvieto, Antonio, et al.
Published: (2025)
by: Orvieto, Antonio, et al.
Published: (2025)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
by: Movahedi, Sajad, et al.
Published: (2024)
by: Movahedi, Sajad, et al.
Published: (2024)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
Explaining Grokking in Transformers through the Lens of Inductive Bias
by: Singh, Jaisidh, et al.
Published: (2026)
by: Singh, Jaisidh, et al.
Published: (2026)
Improved state mixing in higher-order and block diagonal linear recurrent networks
by: Dubinin, Igor, et al.
Published: (2026)
by: Dubinin, Igor, et al.
Published: (2026)
When, Where and Why to Average Weights?
by: Ajroldi, Niccolò, et al.
Published: (2025)
by: Ajroldi, Niccolò, et al.
Published: (2025)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
by: Sieber, Jerome, et al.
Published: (2024)
by: Sieber, Jerome, et al.
Published: (2024)
Thinking into the Future: Latent Lookahead Training for Transformers
by: Noci, Lorenzo, et al.
Published: (2026)
by: Noci, Lorenzo, et al.
Published: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Scaling Behavior of Discrete Diffusion Language Models
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
by: Köprücü, Nursena, et al.
Published: (2024)
by: Köprücü, Nursena, et al.
Published: (2024)
On the low-shot transferability of [V]-Mamba
by: Misra, Diganta, et al.
Published: (2024)
by: Misra, Diganta, et al.
Published: (2024)
Towards Understanding Self-Pretraining for Sequence Classification
by: Coser, Omar, et al.
Published: (2026)
by: Coser, Omar, et al.
Published: (2026)
Some Super-approximation Rates of ReLU Neural Networks for Korobov Functions
by: Li, Yuwen, et al.
Published: (2025)
by: Li, Yuwen, et al.
Published: (2025)
On the Convergence of Gradient Descent for Large Learning Rates
by: Crăciun, Alexandru, et al.
Published: (2024)
by: Crăciun, Alexandru, et al.
Published: (2024)
Fixed-Point RNNs: Interpolating from Diagonal to Dense
by: Movahedi, Sajad, et al.
Published: (2025)
by: Movahedi, Sajad, et al.
Published: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
by: Elhassan, Fay, et al.
Published: (2025)
by: Elhassan, Fay, et al.
Published: (2025)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Generalized Interpolating Discrete Diffusion
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
Learning Consistent Causal Abstraction Networks
by: D'Acunto, Gabriele, et al.
Published: (2026)
by: D'Acunto, Gabriele, et al.
Published: (2026)
Muown: Row-Norm Control for Muon Optimization
by: Lion, Kai, et al.
Published: (2026)
by: Lion, Kai, et al.
Published: (2026)
Learning Sparse Compositional Functions with Norm-Constrained Neural Networks
by: Huang, Shuo, et al.
Published: (2026)
by: Huang, Shuo, et al.
Published: (2026)
Similar Items
-
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
by: Belloni, Annalisa, et al.
Published: (2026) -
The Optimization Landscape of SGD Across the Feature Learning Strength
by: Atanasov, Alexander, et al.
Published: (2024) -
Understanding and Minimising Outlier Features in Neural Network Training
by: He, Bobby, et al.
Published: (2024) -
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)