Variational Learning Finds Flatter Solutions at the Edge of Stability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghosh, Avrajit, Cong, Bai, Yokota, Rio, Ravishankar, Saiprasad, Wang, Rongrong, Tao, Molei, Khan, Mohammad Emtiyaz, Möllenhoff, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Dynamics of Deep Linear Networks Beyond the Edge of Stability
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2025)
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2025)
Improving LoRA with Variational Learning
von: Cong, Bai, et al.
Veröffentlicht: (2025)
von: Cong, Bai, et al.
Veröffentlicht: (2025)
Variational Low-Rank Adaptation Using IVON
von: Cong, Bai, et al.
Veröffentlicht: (2024)
von: Cong, Bai, et al.
Veröffentlicht: (2024)
Optimal Eye Surgeon: Finding Image Priors through Sparse Generators at Initialization
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2024)
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2024)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
von: Kumar, Navish, et al.
Veröffentlicht: (2025)
von: Kumar, Navish, et al.
Veröffentlicht: (2025)
Pruning Unrolled Networks (PUN) at Initialization for MRI Reconstruction Improves Generalization
von: Liang, Shijun, et al.
Veröffentlicht: (2024)
von: Liang, Shijun, et al.
Veröffentlicht: (2024)
Variational Learning is Effective for Large Deep Networks
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
Natural Variational Annealing for Multimodal Optimization
von: LeMinh, Tâm, et al.
Veröffentlicht: (2025)
von: LeMinh, Tâm, et al.
Veröffentlicht: (2025)
Information Geometry of Variational Bayes
von: Khan, Mohammad Emtiyaz
Veröffentlicht: (2025)
von: Khan, Mohammad Emtiyaz
Veröffentlicht: (2025)
Federated ADMM from Bayesian Duality
von: Möllenhoff, Thomas, et al.
Veröffentlicht: (2025)
von: Möllenhoff, Thomas, et al.
Veröffentlicht: (2025)
SVRG and Beyond via Posterior Correction
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
Conformal Prediction via Regression-as-Classification
von: Guha, Etash, et al.
Veröffentlicht: (2024)
von: Guha, Etash, et al.
Veröffentlicht: (2024)
The Memory Perturbation Equation: Understanding Model's Sensitivity to Data
von: Nickl, Peter, et al.
Veröffentlicht: (2023)
von: Nickl, Peter, et al.
Veröffentlicht: (2023)
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
von: Nishida, Keigo, et al.
Veröffentlicht: (2025)
von: Nishida, Keigo, et al.
Veröffentlicht: (2025)
Joint Model and Data Sparsification via the Marginal Likelihood
von: Timans, Alexander, et al.
Veröffentlicht: (2026)
von: Timans, Alexander, et al.
Veröffentlicht: (2026)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
von: Maldonado, Hugo Monzón, et al.
Veröffentlicht: (2024)
von: Maldonado, Hugo Monzón, et al.
Veröffentlicht: (2024)
Model Merging by Uncertainty-Based Gradient Matching
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
The Bayesian Learning Rule
von: Khan, Mohammad Emtiyaz, et al.
Veröffentlicht: (2021)
von: Khan, Mohammad Emtiyaz, et al.
Veröffentlicht: (2021)
Knowledge Adaptation as Posterior Correction
von: Khan, Mohammad Emtiyaz
Veröffentlicht: (2025)
von: Khan, Mohammad Emtiyaz
Veröffentlicht: (2025)
Compact Memory for Continual Logistic Regression
von: Jung, Yohan, et al.
Veröffentlicht: (2025)
von: Jung, Yohan, et al.
Veröffentlicht: (2025)
Improving Generalization of Complex Models under Unbounded Loss Using PAC-Bayes Bounds
von: Zhang, Xitong, et al.
Veröffentlicht: (2023)
von: Zhang, Xitong, et al.
Veröffentlicht: (2023)
Variational Learning Induces Adaptive Label Smoothing
von: Yang, Sin-Han, et al.
Veröffentlicht: (2025)
von: Yang, Sin-Han, et al.
Veröffentlicht: (2025)
Understanding Untrained Deep Models for Inverse Problems: Algorithms and Theory
von: Alkhouri, Ismail, et al.
Veröffentlicht: (2025)
von: Alkhouri, Ismail, et al.
Veröffentlicht: (2025)
Adaptive Local Neighborhood-based Neural Networks for MR Image Reconstruction from Undersampled Data
von: Liang, Shijun, et al.
Veröffentlicht: (2022)
von: Liang, Shijun, et al.
Veröffentlicht: (2022)
Stein's Lemma for the Reparameterization Trick with Exponential Family Mixtures
von: Lin, Wu, et al.
Veröffentlicht: (2019)
von: Lin, Wu, et al.
Veröffentlicht: (2019)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2023)
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2023)
A Dataless Reinforcement Learning Approach to Rounding Hyperplane Optimization for Max-Cut
von: Maliakal, Gabriel, et al.
Veröffentlicht: (2025)
von: Maliakal, Gabriel, et al.
Veröffentlicht: (2025)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
Connecting Federated ADMM to Bayes
von: Swaroop, Siddharth, et al.
Veröffentlicht: (2025)
von: Swaroop, Siddharth, et al.
Veröffentlicht: (2025)
Variational Schrödinger Momentum Diffusion
von: Rojas, Kevin, et al.
Veröffentlicht: (2025)
von: Rojas, Kevin, et al.
Veröffentlicht: (2025)
Learning Gradient-based Mixup with Extrapolation toward Flatter Minima for Domain Generalization
von: Peng, Danni, et al.
Veröffentlicht: (2022)
von: Peng, Danni, et al.
Veröffentlicht: (2022)
Tada-DIP: Input-adaptive Deep Image Prior for One-shot 3D Image Reconstruction
von: Bell, Evan, et al.
Veröffentlicht: (2025)
von: Bell, Evan, et al.
Veröffentlicht: (2025)
Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning
von: Lin, Wu, et al.
Veröffentlicht: (2023)
von: Lin, Wu, et al.
Veröffentlicht: (2023)
A Stein Identity for q-Gaussians with Bounded Support
von: Sklaviadis, Sophia, et al.
Veröffentlicht: (2026)
von: Sklaviadis, Sophia, et al.
Veröffentlicht: (2026)
SITCOM: Step-wise Triple-Consistent Diffusion Sampling for Inverse Problems
von: Alkhouri, Ismail, et al.
Veröffentlicht: (2024)
von: Alkhouri, Ismail, et al.
Veröffentlicht: (2024)
Hard labels sampled from sparse targets mislead rotation invariant algorithms
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2026)
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2026)
DeepTTV: Deep Learning Prediction of Hidden Exoplanet From Transit Timing Variations
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Bilateral Sharpness-Aware Minimization for Flatter Minima
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
Bridging the Gap Between Target Networks and Functional Regularization
von: Piche, Alexandre, et al.
Veröffentlicht: (2022)
von: Piche, Alexandre, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Learning Dynamics of Deep Linear Networks Beyond the Edge of Stability
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2025) -
Improving LoRA with Variational Learning
von: Cong, Bai, et al.
Veröffentlicht: (2025) -
Variational Low-Rank Adaptation Using IVON
von: Cong, Bai, et al.
Veröffentlicht: (2024) -
Optimal Eye Surgeon: Finding Image Priors through Sparse Generators at Initialization
von: Ghosh, Avrajit, et al.
Veröffentlicht: (2024) -
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
von: Kumar, Navish, et al.
Veröffentlicht: (2025)