Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Ze, Zhang, Jian, Wang, Yisen, Qi, Lei, Shi, Yinghuan, Gao, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards the Connection between Activation Sparsity and Flat Minima
by: Peng, Ze, et al.
Published: (2026)
by: Peng, Ze, et al.
Published: (2026)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025)
by: Peng, Ze, et al.
Published: (2025)
Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain Generalization
by: Wang, Xiran, et al.
Published: (2025)
by: Wang, Xiran, et al.
Published: (2025)
An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
Constructing and Exploring Intermediate Domains in Mixed Domain Semi-supervised Medical Image Segmentation
by: Ma, Qinghe, et al.
Published: (2024)
by: Ma, Qinghe, et al.
Published: (2024)
On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
by: Lampert, Christoph H., et al.
Published: (2026)
by: Lampert, Christoph H., et al.
Published: (2026)
Information-Theoretic Generalization Bounds of Replay-based Continual Learning
by: Wen, Wen, et al.
Published: (2025)
by: Wen, Wen, et al.
Published: (2025)
Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRA
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
Generalization and Optimization of SGD with Lookahead
by: Li, Kangcheng, et al.
Published: (2025)
by: Li, Kangcheng, et al.
Published: (2025)
Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning
by: Barnes, L. P., et al.
Published: (2022)
by: Barnes, L. P., et al.
Published: (2022)
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
by: Li, Yayuan, et al.
Published: (2026)
by: Li, Yayuan, et al.
Published: (2026)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
by: Duan, Yue, et al.
Published: (2025)
by: Duan, Yue, et al.
Published: (2025)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026)
by: Tao, Hongyi, et al.
Published: (2026)
MAGIC: Achieving Superior Model Merging via Magnitude Calibration
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
Information Theoretic Lower Bounds for Information Theoretic Upper Bounds
by: Livni, Roi
Published: (2023)
by: Livni, Roi
Published: (2023)
Improved Stability and Generalization Guarantees of the Decentralized SGD Algorithm
by: Bars, Batiste Le, et al.
Published: (2023)
by: Bars, Batiste Le, et al.
Published: (2023)
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
An Inclusive Theoretical Framework of Robust Supervised Contrastive Loss against Label Noise
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
Error Bounds of Supervised Classification from Information-Theoretic Perspective
by: Qi, Binchuan
Published: (2024)
by: Qi, Binchuan
Published: (2024)
Information-Theoretic Generalization Bounds for Sequential Decision Making
by: Futami, Futoshi, et al.
Published: (2026)
by: Futami, Futoshi, et al.
Published: (2026)
Two Facets of SDE Under an Information-Theoretic Lens: Generalization of SGD via Training Trajectories and via Terminal States
by: Wang, Ziqiao, et al.
Published: (2022)
by: Wang, Ziqiao, et al.
Published: (2022)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
by: Jin, Richeng, et al.
Published: (2020)
by: Jin, Richeng, et al.
Published: (2020)
An Augmentation Overlap Theory of Contrastive Learning
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Projection Head is Secretly an Information Bottleneck
by: Ouyang, Zhuo, et al.
Published: (2025)
by: Ouyang, Zhuo, et al.
Published: (2025)
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
by: Tang, Huayi, et al.
Published: (2023)
by: Tang, Huayi, et al.
Published: (2023)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025)
by: Zhang, Yi-Ge, et al.
Published: (2025)
Discrete Curvature Graph Information Bottleneck
by: Fu, Xingcheng, et al.
Published: (2024)
by: Fu, Xingcheng, et al.
Published: (2024)
Empirical Risk Minimization with Shuffled SGD: A Primal-Dual Perspective and Improved Bounds
by: Cai, Xufeng, et al.
Published: (2023)
by: Cai, Xufeng, et al.
Published: (2023)
Information-Theoretic Generalization Bounds for Deep Neural Networks
by: He, Haiyun, et al.
Published: (2024)
by: He, Haiyun, et al.
Published: (2024)
Improving Generalization with Flat Hilbert Bayesian Inference
by: Truong, Tuan, et al.
Published: (2024)
by: Truong, Tuan, et al.
Published: (2024)
Stability and Generalization for Decentralized Markov SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
Topology-aware Generalization of Decentralized SGD
by: Zhu, Tongtian, et al.
Published: (2022)
by: Zhu, Tongtian, et al.
Published: (2022)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
by: Sha, Haichao, et al.
Published: (2023)
by: Sha, Haichao, et al.
Published: (2023)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Accumulative SGD Influence Estimation for Data Attribution
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Information-Theoretic Generalization Bounds for Stochastic Gradient Descent with Predictable Virtual Noise
by: Partohaghighi, Mohammad
Published: (2026)
by: Partohaghighi, Mohammad
Published: (2026)
Unveiling High-Probability Generalization in Decentralized SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
Similar Items
-
Towards the Connection between Activation Sparsity and Flat Minima
by: Peng, Ze, et al.
Published: (2026) -
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025) -
Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain Generalization
by: Wang, Xiran, et al.
Published: (2025) -
An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
by: Duan, Yue, et al.
Published: (2025) -
Constructing and Exploring Intermediate Domains in Mixed Domain Semi-supervised Medical Image Segmentation
by: Ma, Qinghe, et al.
Published: (2024)