Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More
Fuente:
arXiv
Saved in:
| Main Authors: | Yoo, Geonhui, Song, Minhak, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025)
by: Cho, Yeseul, et al.
Published: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
by: Moon, Chaewon, et al.
Published: (2026)
by: Moon, Chaewon, et al.
Published: (2026)
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
by: Shin, Baekrok, et al.
Published: (2026)
by: Shin, Baekrok, et al.
Published: (2026)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
by: Chae, Jiseok, et al.
Published: (2024)
by: Chae, Jiseok, et al.
Published: (2024)
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
A Minimalist Example of Edge-of-Stability and Progressive Sharpening
by: Liu, Liming, et al.
Published: (2025)
by: Liu, Liming, et al.
Published: (2025)
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)
by: Oh, Junsoo, et al.
Published: (2024)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
by: Shin, Baekrok, et al.
Published: (2024)
by: Shin, Baekrok, et al.
Published: (2024)
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo
by: Kim, Hyunsu, et al.
Published: (2025)
by: Kim, Hyunsu, et al.
Published: (2025)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
by: Shi, Ruizhe, et al.
Published: (2025)
by: Shi, Ruizhe, et al.
Published: (2025)
Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Fundamental Benefit of Alternating Updates in Minimax Optimization
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
AcTTA: Rethinking Test-Time Adaptation via Dynamic Activation
by: Kim, Hyeongyu, et al.
Published: (2026)
by: Kim, Hyeongyu, et al.
Published: (2026)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
Uniform Spectral Growth and Convergence of Muon in LoRA-Style Matrix Factorization
by: Kang, Changmin, et al.
Published: (2026)
by: Kang, Changmin, et al.
Published: (2026)
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025)
by: Wang, Kaizheng
Published: (2025)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
by: Ethayarajh, Kawin, et al.
Published: (2021)
by: Ethayarajh, Kawin, et al.
Published: (2021)
Ordinality in Discrete-level Question Difficulty Estimation: Introducing Balanced DRPS and OrderedLogitNN
by: Thuy, Arthur, et al.
Published: (2025)
by: Thuy, Arthur, et al.
Published: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
by: Clara, Gabriel, et al.
Published: (2025)
by: Clara, Gabriel, et al.
Published: (2025)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Leveraging Stochastic Depth Training for Adaptive Inference
by: Korol, Guilherme, et al.
Published: (2025)
by: Korol, Guilherme, et al.
Published: (2025)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
A Minimalist Prompt for Zero-Shot Policy Learning
by: Song, Meng, et al.
Published: (2024)
by: Song, Meng, et al.
Published: (2024)
It's My Data Too: Private ML for Datasets with Multi-User Training Examples
by: Ganesh, Arun, et al.
Published: (2025)
by: Ganesh, Arun, et al.
Published: (2025)
Buffer layers for Test-Time Adaptation
by: Kim, Hyeongyu, et al.
Published: (2025)
by: Kim, Hyeongyu, et al.
Published: (2025)
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
by: Jung, Jeesu, et al.
Published: (2025)
by: Jung, Jeesu, et al.
Published: (2025)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
UnifiedNN: Efficient Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024)
by: Taki, Sifat Ut, et al.
Published: (2024)
Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
Modeling Personalized Difficulty of Rehabilitation Exercises Using Causal Trees
by: Dennler, Nathaniel, et al.
Published: (2025)
by: Dennler, Nathaniel, et al.
Published: (2025)
Understanding the Difficulty of Solving Cauchy Problems with PINNs
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Spectral-factorized Positive-definite Curvature Learning for NN Training
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Similar Items
-
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025) -
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024) -
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025) -
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025) -
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
by: Moon, Chaewon, et al.
Published: (2026)