Dynamics of Meta-learning Representation in the Teacher-student Scenario
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Hui, Yip, Cho Tung, Li, Bo |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Symmetric Perceptron: a Teacher-Student Scenario
par: Catania, Giovanni, et autres
Publié: (2026)
par: Catania, Giovanni, et autres
Publié: (2026)
Applying statistical learning theory to deep learning
par: Gerbelot, Cédric, et autres
Publié: (2023)
par: Gerbelot, Cédric, et autres
Publié: (2023)
Formation of Representations in Neural Networks
par: Ziyin, Liu, et autres
Publié: (2024)
par: Ziyin, Liu, et autres
Publié: (2024)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
par: Thériault, Robin, et autres
Publié: (2024)
par: Thériault, Robin, et autres
Publié: (2024)
Diffusion Operator Geometry of Feedforward Representations
par: Reddy, Kanishka
Publié: (2026)
par: Reddy, Kanishka
Publié: (2026)
Machine learning for sustainable geoenergy: uncertainty, physics and decision-ready inference
par: Menke, Hannah P., et autres
Publié: (2026)
par: Menke, Hannah P., et autres
Publié: (2026)
Dynamical Regimes of Multimodal Diffusion Models
par: Albrychiewicz, Emil, et autres
Publié: (2026)
par: Albrychiewicz, Emil, et autres
Publié: (2026)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
par: Meng, Lineghuan, et autres
Publié: (2024)
par: Meng, Lineghuan, et autres
Publié: (2024)
Asymptotic theory of in-context learning by linear attention
par: Lu, Yue M., et autres
Publié: (2024)
par: Lu, Yue M., et autres
Publié: (2024)
High-dimensional learning of narrow neural networks
par: Cui, Hugo
Publié: (2024)
par: Cui, Hugo
Publié: (2024)
A unified theory of feature learning in RNNs and DNNs
par: Bauer, Jan P., et autres
Publié: (2026)
par: Bauer, Jan P., et autres
Publié: (2026)
A solvable model of learning generative diffusion: theory and insights
par: Cui, Hugo, et autres
Publié: (2025)
par: Cui, Hugo, et autres
Publié: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
par: Cagnetta, Francesco, et autres
Publié: (2025)
par: Cagnetta, Francesco, et autres
Publié: (2025)
Hopfield model with planted patterns: a teacher-student self-supervised learning model
par: Alemanno, Francesco, et autres
Publié: (2023)
par: Alemanno, Francesco, et autres
Publié: (2023)
Asymptotics of feature learning in two-layer networks after one gradient-step
par: Cui, Hugo, et autres
Publié: (2024)
par: Cui, Hugo, et autres
Publié: (2024)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
par: Keup, Christian, et autres
Publié: (2024)
par: Keup, Christian, et autres
Publié: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
par: Lauditi, Clarissa, et autres
Publié: (2025)
par: Lauditi, Clarissa, et autres
Publié: (2025)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
par: Nishiyama, Sota, et autres
Publié: (2025)
par: Nishiyama, Sota, et autres
Publié: (2025)
Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
par: Yang, Ning, et autres
Publié: (2026)
par: Yang, Ning, et autres
Publié: (2026)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
par: D'Amico, Francesco, et autres
Publié: (2025)
par: D'Amico, Francesco, et autres
Publié: (2025)
Infinite Limits of Multi-head Transformer Dynamics
par: Bordelon, Blake, et autres
Publié: (2024)
par: Bordelon, Blake, et autres
Publié: (2024)
A Dynamical Model of Neural Scaling Laws
par: Bordelon, Blake, et autres
Publié: (2024)
par: Bordelon, Blake, et autres
Publié: (2024)
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
par: Cowsik, Aditya, et autres
Publié: (2024)
par: Cowsik, Aditya, et autres
Publié: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
par: Kumar, Tanishq, et autres
Publié: (2023)
par: Kumar, Tanishq, et autres
Publié: (2023)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
par: Troiani, Emanuele, et autres
Publié: (2025)
par: Troiani, Emanuele, et autres
Publié: (2025)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
par: Jain, Anchit, et autres
Publié: (2024)
par: Jain, Anchit, et autres
Publié: (2024)
Dynamical Mean-Field Theory of Self-Attention Neural Networks
par: Poc-López, Ángel, et autres
Publié: (2024)
par: Poc-López, Ángel, et autres
Publié: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
par: Patel, Nishil, et autres
Publié: (2023)
par: Patel, Nishil, et autres
Publié: (2023)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
par: Nicoletti, Flavio, et autres
Publié: (2026)
par: Nicoletti, Flavio, et autres
Publié: (2026)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
par: Radhakrishnan, Anil, et autres
Publié: (2025)
par: Radhakrishnan, Anil, et autres
Publié: (2025)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
par: Montanari, Andrea, et autres
Publié: (2025)
par: Montanari, Andrea, et autres
Publié: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
par: Bordelon, Blake, et autres
Publié: (2026)
par: Bordelon, Blake, et autres
Publié: (2026)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
par: Zambon, Alessandro, et autres
Publié: (2026)
par: Zambon, Alessandro, et autres
Publié: (2026)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
par: Atanasov, Alexander, et autres
Publié: (2025)
par: Atanasov, Alexander, et autres
Publié: (2025)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
par: Mendes, Vicente Conde, et autres
Publié: (2026)
par: Mendes, Vicente Conde, et autres
Publié: (2026)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
par: Bonnaire, Tony, et autres
Publié: (2025)
par: Bonnaire, Tony, et autres
Publié: (2025)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
par: Veiga, Rodrigo, et autres
Publié: (2024)
par: Veiga, Rodrigo, et autres
Publié: (2024)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
par: Nishiyama, Sota, et autres
Publié: (2026)
par: Nishiyama, Sota, et autres
Publié: (2026)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
par: Mainali, Nischal, et autres
Publié: (2025)
par: Mainali, Nischal, et autres
Publié: (2025)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
par: Bordelon, Blake, et autres
Publié: (2025)
par: Bordelon, Blake, et autres
Publié: (2025)
Documents similaires
-
The Symmetric Perceptron: a Teacher-Student Scenario
par: Catania, Giovanni, et autres
Publié: (2026) -
Applying statistical learning theory to deep learning
par: Gerbelot, Cédric, et autres
Publié: (2023) -
Formation of Representations in Neural Networks
par: Ziyin, Liu, et autres
Publié: (2024) -
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
par: Thériault, Robin, et autres
Publié: (2024) -
Diffusion Operator Geometry of Feedforward Representations
par: Reddy, Kanishka
Publié: (2026)