Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Mohua, Beneventano, Pierfrancesco, Dey, Shibshankar, McKinkey, Gareth H., Poggio, Tomaso |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
Does Weight Decay Enhance Training Stability?
by: Saether, Marius, et al.
Published: (2026)
by: Saether, Marius, et al.
Published: (2026)
Too Sharp, Too Sure: When Calibration Follows Curvature
by: Morosini, Alessandro, et al.
Published: (2026)
by: Morosini, Alessandro, et al.
Published: (2026)
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
by: Andreyev, Arseniy, et al.
Published: (2026)
by: Andreyev, Arseniy, et al.
Published: (2026)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
A Cutting-plane and Benders' Decomposition Algorithm for Two-Stage Distributionally Robust Convex programs
by: Luo, Fengqiao, et al.
Published: (2021)
by: Luo, Fengqiao, et al.
Published: (2021)
On Solving Chance-Constrained Models with Gaussian Mixture Distribution
by: Dey, Shibshankar, et al.
Published: (2025)
by: Dey, Shibshankar, et al.
Published: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
Optimization Modeling for Pandemic Vaccine Supply Chain Management: A Review and Future Research Opportunities
by: Dey, Shibshankar, et al.
Published: (2023)
by: Dey, Shibshankar, et al.
Published: (2023)
pAI/MSc: ML Theory Research with Humans on the Loop
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
by: Abdelmoneum, Mahmoud, et al.
Published: (2026)
Iterative regularization in classification via hinge loss diagonal descent
by: Apidopoulos, Vassilis, et al.
Published: (2022)
by: Apidopoulos, Vassilis, et al.
Published: (2022)
When to Forget? Complexity Trade-offs in Machine Unlearning
by: Van Waerebeke, Martin, et al.
Published: (2025)
by: Van Waerebeke, Martin, et al.
Published: (2025)
Geometric Foundations of Tuning without Forgetting in Neural ODEs
by: Bayram, Erkan, et al.
Published: (2025)
by: Bayram, Erkan, et al.
Published: (2025)
Mitigating Forgetting in Continual Learning with Selective Gradient Projection
by: Singh, Anika, et al.
Published: (2026)
by: Singh, Anika, et al.
Published: (2026)
Variance-Reduced $(\varepsilon,δ)-$Unlearning using Forget Set Gradients
by: Van Waerebeke, Martin, et al.
Published: (2026)
by: Van Waerebeke, Martin, et al.
Published: (2026)
I2E2S2R Rumor Spreading Model in Homogeneous Network with Hesitating and Forgetting Mechanisms
by: Hasan, Md. Nahid, et al.
Published: (2025)
by: Hasan, Md. Nahid, et al.
Published: (2025)
Symplectic Inductive Bias for Data-Driven Target Reachability in Hamiltonian Systems
by: Ouyang, Zhuo, et al.
Published: (2026)
by: Ouyang, Zhuo, et al.
Published: (2026)
CoNeT-GIANT: A compressed Newton-type fully distributed optimization algorithm
by: Das, Souvik, et al.
Published: (2025)
by: Das, Souvik, et al.
Published: (2025)
On Convergence Analysis of Network-GIANT: An approximate Hessian-based fully distributed optimization algorithm
by: Das, Souvik, et al.
Published: (2026)
by: Das, Souvik, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
RLS Framework with Segmentation of the Forgetting Profile and Low Rank Updates
by: Stotsky, Alexander
Published: (2025)
by: Stotsky, Alexander
Published: (2025)
On the detection of the presence of malicious components in cyber-physical systems in the almost sure sense
by: Das, Souvik, et al.
Published: (2023)
by: Das, Souvik, et al.
Published: (2023)
Identification of Nonlinear Acyclic Networks in Continuous Time from Nonzero Initial Conditions and Full Excitations
by: Anantharaman, Ramachandran, et al.
Published: (2026)
by: Anantharaman, Ramachandran, et al.
Published: (2026)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
Fast Last-Iterate Convergence of Learning in Games Requires Forgetful Algorithms
by: Cai, Yang, et al.
Published: (2024)
by: Cai, Yang, et al.
Published: (2024)
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Certified Inductive Synthesis for Online Mixed-Integer Optimization
by: Zamponi, Marco, et al.
Published: (2025)
by: Zamponi, Marco, et al.
Published: (2025)
Agentic Systems as Boosting Weak Reasoning Models
by: Sunkaraneni, Varun, et al.
Published: (2026)
by: Sunkaraneni, Varun, et al.
Published: (2026)
Infrastructure Planning for Inductive Charging in Electrified Shuttle Systems
by: Bischoff, Paul, et al.
Published: (2025)
by: Bischoff, Paul, et al.
Published: (2025)
Kaczmarz Projection Algorithms in Moving Window: Performance Improvement via Extended Orthogonality & Forgetting
by: Stotsky, Alexander
Published: (2024)
by: Stotsky, Alexander
Published: (2024)
An Initial Condition-Dependent Neural Network Approach for Optimal Control Problems
by: Rubel, Mominul, et al.
Published: (2025)
by: Rubel, Mominul, et al.
Published: (2025)
$k$-Inductive and Interpolation-Inspired Barrier Certificates for Stochastic Dynamical Systems
by: Oumer, Mohammed Adib, et al.
Published: (2025)
by: Oumer, Mohammed Adib, et al.
Published: (2025)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
HBNET-GIANT: A communication-efficient accelerated Newton-type fully distributed optimization algorithm
by: Das, Souvik, et al.
Published: (2025)
by: Das, Souvik, et al.
Published: (2025)
Optimizing rake-links independently of timetables in railway operations
by: Dey, Sourav
Published: (2025)
by: Dey, Sourav
Published: (2025)
Nonlinear Optimal Guidance for Impact Time Control with Field-of-View Constraint
by: Lu, Fangmin, et al.
Published: (2025)
by: Lu, Fangmin, et al.
Published: (2025)
Similar Items
-
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024) -
Does Weight Decay Enhance Training Stability?
by: Saether, Marius, et al.
Published: (2026) -
Too Sharp, Too Sure: When Calibration Follows Curvature
by: Morosini, Alessandro, et al.
Published: (2026) -
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
by: Andreyev, Arseniy, et al.
Published: (2026) -
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)