Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Shenyang, Liao, Boyao, Ouyang, Zhuoli, Pang, Tianyu, Song, Minhak, Yang, Yaoqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
KCES: Training-Free Defense for Robust Graph Neural Networks via Kernel Complexity
by: Jia, Yaning, et al.
Published: (2025)
by: Jia, Yaning, et al.
Published: (2025)
Model Balancing Helps Low-data Training and Fine-tuning
by: Liu, Zihang, et al.
Published: (2024)
by: Liu, Zihang, et al.
Published: (2024)
Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More
by: Yoo, Geonhui, et al.
Published: (2025)
by: Yoo, Geonhui, et al.
Published: (2025)
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
by: Liu, Zihang, et al.
Published: (2025)
by: Liu, Zihang, et al.
Published: (2025)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
by: Dereziński, Michał, et al.
Published: (2026)
by: Dereziński, Michał, et al.
Published: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Graded Suspiciousness of Adversarial Texts to Human
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
by: Tonni, Shakila Mahjabin, et al.
Published: (2024)
High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
by: Jagannath, Aukosh, et al.
Published: (2025)
by: Jagannath, Aukosh, et al.
Published: (2025)
Structured and Fast Optimization: The Kronecker SGD Algorithm
by: Song, Zhao, et al.
Published: (2023)
by: Song, Zhao, et al.
Published: (2023)
Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer Learning
by: Deng, Yuyang, et al.
Published: (2025)
by: Deng, Yuyang, et al.
Published: (2025)
Machine Learning-Based Detection and Analysis of Suspicious Activities in Bitcoin Wallet Transactions in the USA
by: Islam, Md Zahidul, et al.
Published: (2025)
by: Islam, Md Zahidul, et al.
Published: (2025)
Understanding Generalization from Embedding Dimension and Distributional Convergence
by: Yu, Junjie, et al.
Published: (2026)
by: Yu, Junjie, et al.
Published: (2026)
Data Selection for LLM Alignment Using Fine-Grained Preferences
by: Zhang, Jia, et al.
Published: (2025)
by: Zhang, Jia, et al.
Published: (2025)
Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning
by: Cui, Peng, et al.
Published: (2026)
by: Cui, Peng, et al.
Published: (2026)
Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
by: Hu, Yuanzhe, et al.
Published: (2025)
by: Hu, Yuanzhe, et al.
Published: (2025)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
by: Shi, Ruizhe, et al.
Published: (2025)
by: Shi, Ruizhe, et al.
Published: (2025)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
by: Song, Yuhang, et al.
Published: (2024)
by: Song, Yuhang, et al.
Published: (2024)
Faster Game Solving via Asymmetry of Step Sizes
by: Meng, Linjian, et al.
Published: (2025)
by: Meng, Linjian, et al.
Published: (2025)
Convergence Analysis of SGD under Expected Smoothness
by: Kawamoto, Yuta, et al.
Published: (2025)
by: Kawamoto, Yuta, et al.
Published: (2025)
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
by: Cui, Peng, et al.
Published: (2026)
by: Cui, Peng, et al.
Published: (2026)
Topology-aware Generalization of Decentralized SGD
by: Zhu, Tongtian, et al.
Published: (2022)
by: Zhu, Tongtian, et al.
Published: (2022)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
Tight Analysis of Decentralized SGD: A Markov Chain Perspective
by: Versini, Lucas, et al.
Published: (2026)
by: Versini, Lucas, et al.
Published: (2026)
A Three-regime Model of Network Pruning
by: Zhou, Yefan, et al.
Published: (2023)
by: Zhou, Yefan, et al.
Published: (2023)
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
by: Horii, Hiroshi, et al.
Published: (2025)
by: Horii, Hiroshi, et al.
Published: (2025)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Exact Mean Square Linear Stability Analysis for SGD
by: Mulayoff, Rotem, et al.
Published: (2023)
by: Mulayoff, Rotem, et al.
Published: (2023)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
Similar Items
-
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026) -
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
by: Deng, Shenyang, et al.
Published: (2026) -
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026) -
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024) -
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)