Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Nishida, Keigo, Kıral, Eren Mehmet, Bannai, Kenichi, Khan, Mohammad Emtiyaz, Möllenhoff, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Samples to Probe Trained Models
by: Kıral, Eren Mehmet, et al.
Published: (2025)
by: Kıral, Eren Mehmet, et al.
Published: (2025)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)
by: Kumar, Navish, et al.
Published: (2025)
Federated ADMM from Bayesian Duality
by: Möllenhoff, Thomas, et al.
Published: (2025)
by: Möllenhoff, Thomas, et al.
Published: (2025)
SVRG and Beyond via Posterior Correction
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Conformal Prediction via Regression-as-Classification
by: Guha, Etash, et al.
Published: (2024)
by: Guha, Etash, et al.
Published: (2024)
The Memory Perturbation Equation: Understanding Model's Sensitivity to Data
by: Nickl, Peter, et al.
Published: (2023)
by: Nickl, Peter, et al.
Published: (2023)
Joint Model and Data Sparsification via the Marginal Likelihood
by: Timans, Alexander, et al.
Published: (2026)
by: Timans, Alexander, et al.
Published: (2026)
Natural Variational Annealing for Multimodal Optimization
by: LeMinh, Tâm, et al.
Published: (2025)
by: LeMinh, Tâm, et al.
Published: (2025)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
Variational Low-Rank Adaptation Using IVON
by: Cong, Bai, et al.
Published: (2024)
by: Cong, Bai, et al.
Published: (2024)
Model Merging by Uncertainty-Based Gradient Matching
by: Daheim, Nico, et al.
Published: (2023)
by: Daheim, Nico, et al.
Published: (2023)
Knowledge Adaptation as Posterior Correction
by: Khan, Mohammad Emtiyaz
Published: (2025)
by: Khan, Mohammad Emtiyaz
Published: (2025)
Information Geometry of Variational Bayes
by: Khan, Mohammad Emtiyaz
Published: (2025)
by: Khan, Mohammad Emtiyaz
Published: (2025)
Improving LoRA with Variational Learning
by: Cong, Bai, et al.
Published: (2025)
by: Cong, Bai, et al.
Published: (2025)
Compact Memory for Continual Logistic Regression
by: Jung, Yohan, et al.
Published: (2025)
by: Jung, Yohan, et al.
Published: (2025)
The Bayesian Learning Rule
by: Khan, Mohammad Emtiyaz, et al.
Published: (2021)
by: Khan, Mohammad Emtiyaz, et al.
Published: (2021)
Variational Learning Finds Flatter Solutions at the Edge of Stability
by: Ghosh, Avrajit, et al.
Published: (2025)
by: Ghosh, Avrajit, et al.
Published: (2025)
Variational Learning is Effective for Large Deep Networks
by: Shen, Yuesong, et al.
Published: (2024)
by: Shen, Yuesong, et al.
Published: (2024)
Stein's Lemma for the Reparameterization Trick with Exponential Family Mixtures
by: Lin, Wu, et al.
Published: (2019)
by: Lin, Wu, et al.
Published: (2019)
Connecting Federated ADMM to Bayes
by: Swaroop, Siddharth, et al.
Published: (2025)
by: Swaroop, Siddharth, et al.
Published: (2025)
Bridging the Gap Between Target Networks and Functional Regularization
by: Piche, Alexandre, et al.
Published: (2022)
by: Piche, Alexandre, et al.
Published: (2022)
A Stein Identity for q-Gaussians with Bounded Support
by: Sklaviadis, Sophia, et al.
Published: (2026)
by: Sklaviadis, Sophia, et al.
Published: (2026)
Variational Learning Induces Adaptive Label Smoothing
by: Yang, Sin-Han, et al.
Published: (2025)
by: Yang, Sin-Han, et al.
Published: (2025)
Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning
by: Lin, Wu, et al.
Published: (2023)
by: Lin, Wu, et al.
Published: (2023)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
Sample Efficient Robot Learning in Supervised Effect Prediction Tasks
by: Eren, Mehmet Arda, et al.
Published: (2024)
by: Eren, Mehmet Arda, et al.
Published: (2024)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Stable Training of Normalizing Flows for High-dimensional Variational Inference
by: Andrade, Daniel
Published: (2024)
by: Andrade, Daniel
Published: (2024)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
by: Müller, Lorenz K., et al.
Published: (2025)
by: Müller, Lorenz K., et al.
Published: (2025)
Stable-by-Design Neural Network-Based LPV State-Space Models for System Identification
by: Sertbaş, Ahmet Eren, et al.
Published: (2025)
by: Sertbaş, Ahmet Eren, et al.
Published: (2025)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
Collage: Light-Weight Low-Precision Strategy for LLM Training
by: Yu, Tao, et al.
Published: (2024)
by: Yu, Tao, et al.
Published: (2024)
A Logic for Expressing Log-Precision Transformers
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
by: Zhao, Shen-Yi, et al.
Published: (2020)
by: Zhao, Shen-Yi, et al.
Published: (2020)
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Context-Based Echo State Networks with Prediction Confidence for Human-Robot Shared Control
by: Amirshirzad, Negin, et al.
Published: (2024)
by: Amirshirzad, Negin, et al.
Published: (2024)
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Effective Capacitance Modeling Using Graph Neural Networks
by: Dogan, Eren, et al.
Published: (2025)
by: Dogan, Eren, et al.
Published: (2025)
Similar Items
-
Generating Samples to Probe Trained Models
by: Kıral, Eren Mehmet, et al.
Published: (2025) -
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025) -
Federated ADMM from Bayesian Duality
by: Möllenhoff, Thomas, et al.
Published: (2025) -
SVRG and Beyond via Posterior Correction
by: Daheim, Nico, et al.
Published: (2025) -
Conformal Prediction via Regression-as-Classification
by: Guha, Etash, et al.
Published: (2024)