Super Consistency of Neural Network Landscapes and Learning Rate Transfer
Fuente:
arXiv
Salvato in:
| Autori principali: | Noci, Lorenzo, Meterez, Alexandru, Hofmann, Thomas, Orvieto, Antonio |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
di: Belloni, Annalisa, et al.
Pubblicazione: (2026)
di: Belloni, Annalisa, et al.
Pubblicazione: (2026)
The Optimization Landscape of SGD Across the Feature Learning Strength
di: Atanasov, Alexander, et al.
Pubblicazione: (2024)
di: Atanasov, Alexander, et al.
Pubblicazione: (2024)
Understanding and Minimising Outlier Features in Neural Network Training
di: He, Bobby, et al.
Pubblicazione: (2024)
di: He, Bobby, et al.
Pubblicazione: (2024)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
di: Islamov, Rustem, et al.
Pubblicazione: (2024)
di: Islamov, Rustem, et al.
Pubblicazione: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
An Uncertainty Principle for Linear Recurrent Neural Networks
di: François, Alexandre, et al.
Pubblicazione: (2025)
di: François, Alexandre, et al.
Pubblicazione: (2025)
How Good is a Single Basin?
di: Lion, Kai, et al.
Pubblicazione: (2024)
di: Lion, Kai, et al.
Pubblicazione: (2024)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
The Importance of Being Lazy: Scaling Limits of Continual Learning
di: Graldi, Jacopo, et al.
Pubblicazione: (2025)
di: Graldi, Jacopo, et al.
Pubblicazione: (2025)
Recurrent Distance Filtering for Graph Representation Learning
di: Ding, Yuhui, et al.
Pubblicazione: (2023)
di: Ding, Yuhui, et al.
Pubblicazione: (2023)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
di: Zhao, Rosie, et al.
Pubblicazione: (2025)
di: Zhao, Rosie, et al.
Pubblicazione: (2025)
Adam Simplified: Bias Correction Debunked
di: Laing, Sam, et al.
Pubblicazione: (2025)
di: Laing, Sam, et al.
Pubblicazione: (2025)
Revisiting associative recall in modern recurrent models
di: Okpekpe, Destiny, et al.
Pubblicazione: (2025)
di: Okpekpe, Destiny, et al.
Pubblicazione: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
di: Orvieto, Antonio, et al.
Pubblicazione: (2024)
di: Orvieto, Antonio, et al.
Pubblicazione: (2024)
In Search of Adam's Secret Sauce
di: Orvieto, Antonio, et al.
Pubblicazione: (2025)
di: Orvieto, Antonio, et al.
Pubblicazione: (2025)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
di: Movahedi, Sajad, et al.
Pubblicazione: (2024)
di: Movahedi, Sajad, et al.
Pubblicazione: (2024)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
Explaining Grokking in Transformers through the Lens of Inductive Bias
di: Singh, Jaisidh, et al.
Pubblicazione: (2026)
di: Singh, Jaisidh, et al.
Pubblicazione: (2026)
Improved state mixing in higher-order and block diagonal linear recurrent networks
di: Dubinin, Igor, et al.
Pubblicazione: (2026)
di: Dubinin, Igor, et al.
Pubblicazione: (2026)
When, Where and Why to Average Weights?
di: Ajroldi, Niccolò, et al.
Pubblicazione: (2025)
di: Ajroldi, Niccolò, et al.
Pubblicazione: (2025)
Some Super-approximation Rates of ReLU Neural Networks for Korobov Functions
di: Li, Yuwen, et al.
Pubblicazione: (2025)
di: Li, Yuwen, et al.
Pubblicazione: (2025)
Scaling Behavior of Discrete Diffusion Language Models
di: von Rütte, Dimitri, et al.
Pubblicazione: (2025)
di: von Rütte, Dimitri, et al.
Pubblicazione: (2025)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
di: Sieber, Jerome, et al.
Pubblicazione: (2024)
di: Sieber, Jerome, et al.
Pubblicazione: (2024)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
di: Srećković, Teodora, et al.
Pubblicazione: (2025)
di: Srećković, Teodora, et al.
Pubblicazione: (2025)
Thinking into the Future: Latent Lookahead Training for Transformers
di: Noci, Lorenzo, et al.
Pubblicazione: (2026)
di: Noci, Lorenzo, et al.
Pubblicazione: (2026)
On the Convergence of Gradient Descent for Large Learning Rates
di: Crăciun, Alexandru, et al.
Pubblicazione: (2024)
di: Crăciun, Alexandru, et al.
Pubblicazione: (2024)
Towards Understanding Self-Pretraining for Sequence Classification
di: Coser, Omar, et al.
Pubblicazione: (2026)
di: Coser, Omar, et al.
Pubblicazione: (2026)
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
di: Köprücü, Nursena, et al.
Pubblicazione: (2024)
di: Köprücü, Nursena, et al.
Pubblicazione: (2024)
On the low-shot transferability of [V]-Mamba
di: Misra, Diganta, et al.
Pubblicazione: (2024)
di: Misra, Diganta, et al.
Pubblicazione: (2024)
Learning Sparse Compositional Functions with Norm-Constrained Neural Networks
di: Huang, Shuo, et al.
Pubblicazione: (2026)
di: Huang, Shuo, et al.
Pubblicazione: (2026)
Learning Consistent Causal Abstraction Networks
di: D'Acunto, Gabriele, et al.
Pubblicazione: (2026)
di: D'Acunto, Gabriele, et al.
Pubblicazione: (2026)
Realtime-Capable Hybrid Spiking Neural Networks for Neural Decoding of Cortical Activity
di: Krausse, Jann, et al.
Pubblicazione: (2025)
di: Krausse, Jann, et al.
Pubblicazione: (2025)
Fixed-Point RNNs: Interpolating from Diagonal to Dense
di: Movahedi, Sajad, et al.
Pubblicazione: (2025)
di: Movahedi, Sajad, et al.
Pubblicazione: (2025)
Memory-Consistent Neural Networks for Imitation Learning
di: Sridhar, Kaustubh, et al.
Pubblicazione: (2023)
di: Sridhar, Kaustubh, et al.
Pubblicazione: (2023)
Time-Warping Recurrent Neural Networks for Transfer Learning
di: Hirschi, Jonathon
Pubblicazione: (2026)
di: Hirschi, Jonathon
Pubblicazione: (2026)
Visualization and Analysis of the Loss Landscape in Graph Neural Networks
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
di: Belloni, Annalisa, et al.
Pubblicazione: (2026) -
The Optimization Landscape of SGD Across the Feature Learning Strength
di: Atanasov, Alexander, et al.
Pubblicazione: (2024) -
Understanding and Minimising Outlier Features in Neural Network Training
di: He, Bobby, et al.
Pubblicazione: (2024) -
Loss Landscape Characterization of Neural Networks without Over-Parametrization
di: Islamov, Rustem, et al.
Pubblicazione: (2024) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)