To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets
Fuente:
arXiv
Salvato in:
| Autori principali: | Doshi, Darshil, Das, Aritra, He, Tianyu, Gromov, Andrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
di: He, Tianyu, et al.
Pubblicazione: (2024)
di: He, Tianyu, et al.
Pubblicazione: (2024)
Grokking Modular Polynomials
di: Doshi, Darshil, et al.
Pubblicazione: (2024)
di: Doshi, Darshil, et al.
Pubblicazione: (2024)
(How) Can Transformers Predict Pseudo-Random Numbers?
di: Tao, Tao, et al.
Pubblicazione: (2025)
di: Tao, Tao, et al.
Pubblicazione: (2025)
Towards Distributed Neural Architectures
di: Cowsik, Aditya, et al.
Pubblicazione: (2025)
di: Cowsik, Aditya, et al.
Pubblicazione: (2025)
On the origin of neural scaling laws: from random graphs to natural language
di: Barkeshli, Maissam, et al.
Pubblicazione: (2026)
di: Barkeshli, Maissam, et al.
Pubblicazione: (2026)
Differential learning kinetics govern the transition from memorization to generalization during in-context learning
di: Nguyen, Alex, et al.
Pubblicazione: (2024)
di: Nguyen, Alex, et al.
Pubblicazione: (2024)
Generative diffusion for perceptron problems: statistical physics analysis and efficient algorithms
di: Demyanenko, Elizaveta, et al.
Pubblicazione: (2025)
di: Demyanenko, Elizaveta, et al.
Pubblicazione: (2025)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
di: Keup, Christian, et al.
Pubblicazione: (2024)
di: Keup, Christian, et al.
Pubblicazione: (2024)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
Training neural networks with structured noise improves classification and generalization
di: Benedetti, Marco, et al.
Pubblicazione: (2023)
di: Benedetti, Marco, et al.
Pubblicazione: (2023)
Homophily modulates double descent generalization in graph convolution networks
di: Shi, Cheng, et al.
Pubblicazione: (2022)
di: Shi, Cheng, et al.
Pubblicazione: (2022)
Asymptotic generalization error of a single-layer graph convolutional network
di: Duranthon, O., et al.
Pubblicazione: (2024)
di: Duranthon, O., et al.
Pubblicazione: (2024)
Regularization, early-stopping and dreaming: a Hopfield-like setup to address generalization and overfitting
di: Agliari, Elena, et al.
Pubblicazione: (2023)
di: Agliari, Elena, et al.
Pubblicazione: (2023)
Topological Exploration of High-Dimensional Empirical Risk Landscapes: general approach, and applications to phase retrieval
di: Maillard, Antoine, et al.
Pubblicazione: (2026)
di: Maillard, Antoine, et al.
Pubblicazione: (2026)
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2023)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2023)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
di: Mendes, Vicente Conde, et al.
Pubblicazione: (2026)
di: Mendes, Vicente Conde, et al.
Pubblicazione: (2026)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
di: Kalaj, Silvio, et al.
Pubblicazione: (2024)
di: Kalaj, Silvio, et al.
Pubblicazione: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
di: Patel, Nishil, et al.
Pubblicazione: (2023)
di: Patel, Nishil, et al.
Pubblicazione: (2023)
The Quantization Model of Neural Scaling
di: Michaud, Eric J., et al.
Pubblicazione: (2023)
di: Michaud, Eric J., et al.
Pubblicazione: (2023)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
di: Dawid, Anna, et al.
Pubblicazione: (2023)
di: Dawid, Anna, et al.
Pubblicazione: (2023)
How does training shape the Riemannian geometry of neural network representations?
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2023)
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2023)
Grokking as a First Order Phase Transition in Two Layer Networks
di: Rubin, Noa, et al.
Pubblicazione: (2023)
di: Rubin, Noa, et al.
Pubblicazione: (2023)
A universal approximation theorem for nonlinear resistive networks
di: Scellier, Benjamin, et al.
Pubblicazione: (2023)
di: Scellier, Benjamin, et al.
Pubblicazione: (2023)
High-dimensional Asymptotics of Denoising Autoencoders
di: Cui, Hugo, et al.
Pubblicazione: (2023)
di: Cui, Hugo, et al.
Pubblicazione: (2023)
On the different regimes of Stochastic Gradient Descent
di: Sclocchi, Antonio, et al.
Pubblicazione: (2023)
di: Sclocchi, Antonio, et al.
Pubblicazione: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
di: Kumar, Tanishq, et al.
Pubblicazione: (2023)
di: Kumar, Tanishq, et al.
Pubblicazione: (2023)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
di: Francazi, Emanuele, et al.
Pubblicazione: (2023)
di: Francazi, Emanuele, et al.
Pubblicazione: (2023)
Deep neural networks from the perspective of ergodic theory
di: Zhang, Fan
Pubblicazione: (2023)
di: Zhang, Fan
Pubblicazione: (2023)
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
di: Kühn, Marcel, et al.
Pubblicazione: (2023)
di: Kühn, Marcel, et al.
Pubblicazione: (2023)
Field theory for optimal signal propagation in ResNets
di: Fischer, Kirsten, et al.
Pubblicazione: (2023)
di: Fischer, Kirsten, et al.
Pubblicazione: (2023)
DCEM: A deep complementary energy method for solid mechanics
di: Wang, Yizheng, et al.
Pubblicazione: (2023)
di: Wang, Yizheng, et al.
Pubblicazione: (2023)
Applying statistical learning theory to deep learning
di: Gerbelot, Cédric, et al.
Pubblicazione: (2023)
di: Gerbelot, Cédric, et al.
Pubblicazione: (2023)
Finite-time Lyapunov exponents of deep neural networks
di: Storm, L., et al.
Pubblicazione: (2023)
di: Storm, L., et al.
Pubblicazione: (2023)
The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
di: Mao, Jialin, et al.
Pubblicazione: (2023)
di: Mao, Jialin, et al.
Pubblicazione: (2023)
Dataset-Free Weight-Initialization on Restricted Boltzmann Machine
di: Yasuda, Muneki, et al.
Pubblicazione: (2024)
di: Yasuda, Muneki, et al.
Pubblicazione: (2024)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
di: Xu, Yizhou, et al.
Pubblicazione: (2025)
di: Xu, Yizhou, et al.
Pubblicazione: (2025)
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
di: Duranthon, O., et al.
Pubblicazione: (2025)
di: Duranthon, O., et al.
Pubblicazione: (2025)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
di: Sagitova, M., et al.
Pubblicazione: (2026)
di: Sagitova, M., et al.
Pubblicazione: (2026)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
di: Wang, Shanshan, et al.
Pubblicazione: (2025)
di: Wang, Shanshan, et al.
Pubblicazione: (2025)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
di: Ruben, Benjamin S., et al.
Pubblicazione: (2024)
di: Ruben, Benjamin S., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
di: He, Tianyu, et al.
Pubblicazione: (2024) -
Grokking Modular Polynomials
di: Doshi, Darshil, et al.
Pubblicazione: (2024) -
(How) Can Transformers Predict Pseudo-Random Numbers?
di: Tao, Tao, et al.
Pubblicazione: (2025) -
Towards Distributed Neural Architectures
di: Cowsik, Aditya, et al.
Pubblicazione: (2025) -
On the origin of neural scaling laws: from random graphs to natural language
di: Barkeshli, Maissam, et al.
Pubblicazione: (2026)