Lazy vs hasty: linearization in deep networks impacts learning schedule based on example difficulty
Fuente:
arXiv
Saved in:
| Main Authors: | George, Thomas, Lajoie, Guillaume, Baratin, Aristide |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lookbehind-SAM: k steps back, 1 step forward
by: Mordido, Gonçalo, et al.
Published: (2023)
by: Mordido, Gonçalo, et al.
Published: (2023)
Manifold Metric: A Loss Landscape Approach for Predicting Model Performance
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
by: Jain, Anchit, et al.
Published: (2024)
by: Jain, Anchit, et al.
Published: (2024)
Any-Property-Conditional Molecule Generation with Self-Criticism using Spanning Trees
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2024)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2024)
Generating $π$-Functional Molecules Using STGG+ with Active Learning
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2025)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2025)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
How connectivity structure shapes rich and lazy learning in neural circuits
by: Liu, Yuhan Helena, et al.
Published: (2023)
by: Liu, Yuhan Helena, et al.
Published: (2023)
Partially Lazy Gradient Descent for Smoothed Online Learning
by: Mhaisen, Naram, et al.
Published: (2026)
by: Mhaisen, Naram, et al.
Published: (2026)
Bidirectional Information Flow (BIF) -- A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization
by: Guerra, Juan D., et al.
Published: (2025)
by: Guerra, Juan D., et al.
Published: (2025)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
by: Wu, Xinquan, et al.
Published: (2025)
by: Wu, Xinquan, et al.
Published: (2025)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
How deep is your network? Deep vs. shallow learning of transfer operators
by: Tabish, Mohammad, et al.
Published: (2025)
by: Tabish, Mohammad, et al.
Published: (2025)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
by: Dufort-Labbé, Simon, et al.
Published: (2024)
by: Dufort-Labbé, Simon, et al.
Published: (2024)
Mislabeled examples detection viewed as probing machine learning models: concepts, survey and extensive benchmark
by: George, Thomas, et al.
Published: (2024)
by: George, Thomas, et al.
Published: (2024)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Does learning the right latent variables necessarily improve in-context learning?
by: Mittal, Sarthak, et al.
Published: (2024)
by: Mittal, Sarthak, et al.
Published: (2024)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
Celo: Training Versatile Learned Optimizers on a Compute Diet
by: Moudgil, Abhinav, et al.
Published: (2025)
by: Moudgil, Abhinav, et al.
Published: (2025)
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Unsupervised Concept Discovery Mitigates Spurious Correlations
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Calibration improves detection of mislabeled examples
by: Chibane, Ilies, et al.
Published: (2025)
by: Chibane, Ilies, et al.
Published: (2025)
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
by: Williams, Ezekiel, et al.
Published: (2026)
by: Williams, Ezekiel, et al.
Published: (2026)
Solving the flexible job-shop scheduling problem through an enhanced deep reinforcement learning approach
by: Echeverria, Imanol, et al.
Published: (2023)
by: Echeverria, Imanol, et al.
Published: (2023)
Sufficient conditions for offline reactivation in recurrent neural networks
by: Krishna, Nanda H., et al.
Published: (2025)
by: Krishna, Nanda H., et al.
Published: (2025)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
by: Malviya, Pranshu, et al.
Published: (2023)
by: Malviya, Pranshu, et al.
Published: (2023)
In-context learning and Occam's razor
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Interpreting and generalizing deep learning in physics-based problems with functional linear models
by: Arzani, Amirhossein, et al.
Published: (2023)
by: Arzani, Amirhossein, et al.
Published: (2023)
Reinforcement learning-based dynamic cleaning scheduling framework for solar energy system
by: An, Heungjo
Published: (2026)
by: An, Heungjo
Published: (2026)
In value-based deep reinforcement learning, a pruned network is a good network
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Lazy FSCA for Unsupervised Variable Selection
by: Zocco, Federico, et al.
Published: (2021)
by: Zocco, Federico, et al.
Published: (2021)
An invariance constrained deep learning network for PDE discovery
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
The Importance of Being Lazy: Scaling Limits of Continual Learning
by: Graldi, Jacopo, et al.
Published: (2025)
by: Graldi, Jacopo, et al.
Published: (2025)
Decision-focused learning for optimal PV-Battery scheduling
by: Depoortere, Joris, et al.
Published: (2026)
by: Depoortere, Joris, et al.
Published: (2026)
DURENDAL: Graph deep learning framework for temporal heterogeneous networks
by: Dileo, Manuel, et al.
Published: (2023)
by: Dileo, Manuel, et al.
Published: (2023)
Applications of deep reinforcement learning to urban transit network design
by: Holliday, Andrew
Published: (2025)
by: Holliday, Andrew
Published: (2025)
Version age-based client scheduling policy for federated learning
by: Hu, Xinyi, et al.
Published: (2024)
by: Hu, Xinyi, et al.
Published: (2024)
Similar Items
-
Lookbehind-SAM: k steps back, 1 step forward
by: Mordido, Gonçalo, et al.
Published: (2023) -
Manifold Metric: A Loss Landscape Approach for Predicting Model Performance
by: Malviya, Pranshu, et al.
Published: (2024) -
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
by: Jain, Anchit, et al.
Published: (2024) -
Any-Property-Conditional Molecule Generation with Self-Criticism using Spanning Trees
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2024) -
Generating $π$-Functional Molecules Using STGG+ with Active Learning
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2025)