SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sadrtdinov, Ildus, Klimov, Ivan, Lobacheva, Ekaterina, Vetrov, Dmitry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023)
Where Do Large Learning Rates Lead Us?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2024)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2024)
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026)
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026)
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
von: Bartosh, Grigory, et al.
Veröffentlicht: (2025)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2025)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
Neural Diffusion Models
von: Bartosh, Grigory, et al.
Veröffentlicht: (2023)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2023)
Guided Star-Shaped Masked Diffusion
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
von: Kovalev, Dmitry
Veröffentlicht: (2026)
von: Kovalev, Dmitry
Veröffentlicht: (2026)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
Regularized Distribution Matching Distillation for One-step Unpaired Image-to-Image Translation
von: Rakitin, Denis, et al.
Veröffentlicht: (2024)
von: Rakitin, Denis, et al.
Veröffentlicht: (2024)
Generative Flow Networks as Entropy-Regularized RL
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
Direct Bethe Free Energy Minimization for Bayesian Neural Network
von: Prochazka, Pavel
Veröffentlicht: (2026)
von: Prochazka, Pavel
Veröffentlicht: (2026)
Gradual Optimization Learning for Conformational Energy Minimization
von: Tsypin, Artem, et al.
Veröffentlicht: (2023)
von: Tsypin, Artem, et al.
Veröffentlicht: (2023)
Neural Flow Diffusion Models: Learnable Forward Process for Improved Diffusion Modelling
von: Bartosh, Grigory, et al.
Veröffentlicht: (2024)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2024)
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
von: Wan, Yijun, et al.
Veröffentlicht: (2023)
von: Wan, Yijun, et al.
Veröffentlicht: (2023)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
GLGENN: A Novel Parameter-Light Equivariant Neural Networks Architecture Based on Clifford Geometric Algebras
von: Filimoshina, Ekaterina, et al.
Veröffentlicht: (2025)
von: Filimoshina, Ekaterina, et al.
Veröffentlicht: (2025)
Solvation Free Energies from Neural Thermodynamic Integration
von: Máté, Bálint, et al.
Veröffentlicht: (2024)
von: Máté, Bálint, et al.
Veröffentlicht: (2024)
Principled Pruning of Bayesian Neural Networks through Variational Free Energy Minimization
von: Beckers, Jim, et al.
Veröffentlicht: (2022)
von: Beckers, Jim, et al.
Veröffentlicht: (2022)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
von: Kovalev, Dmitry
Veröffentlicht: (2025)
von: Kovalev, Dmitry
Veröffentlicht: (2025)
Neural Thermodynamic Integration: Free Energies from Energy-based Diffusion Models
von: Máté, Bálint, et al.
Veröffentlicht: (2024)
von: Máté, Bálint, et al.
Veröffentlicht: (2024)
From SGD to Spectra: A Theory of Neural Network Weight Dynamics
von: Olsen, Brian Richard, et al.
Veröffentlicht: (2025)
von: Olsen, Brian Richard, et al.
Veröffentlicht: (2025)
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
Making SGD Parameter-Free
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
von: Carmon, Yair, et al.
Veröffentlicht: (2022)
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
von: Vu, Evgeniia, et al.
Veröffentlicht: (2025)
von: Vu, Evgeniia, et al.
Veröffentlicht: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
von: Sun, Ying, et al.
Veröffentlicht: (2024)
von: Sun, Ying, et al.
Veröffentlicht: (2024)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
von: Bethune, Louis, et al.
Veröffentlicht: (2023)
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
von: Horii, Hiroshi, et al.
Veröffentlicht: (2025)
von: Horii, Hiroshi, et al.
Veröffentlicht: (2025)
Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
von: Wang, Puyu, et al.
Veröffentlicht: (2026)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
von: Peng, Hanyang, et al.
Veröffentlicht: (2025)
Neural Network Training via Stochastic Alternating Minimization with Trainable Step Sizes
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
Gradient-Free Training of Quantized Neural Networks
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Improving GFlowNets with Monte Carlo Tree Search
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
Training-Free Cross-Architecture Merging for Graph Neural Networks
von: Bhattacharya, Rishabh, et al.
Veröffentlicht: (2026)
von: Bhattacharya, Rishabh, et al.
Veröffentlicht: (2026)
Asynchronous Local-SGD Training for Language Modeling
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025) -
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023) -
Where Do Large Learning Rates Lead Us?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2024) -
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026) -
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
von: Bartosh, Grigory, et al.
Veröffentlicht: (2025)