Worker Disagreement Reveals Sharp Directions in Local SGD

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dimlioglu, Tolga, Topollai, Kristi, Choromanska, Anna
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916052966834176
author Dimlioglu, Tolga
Topollai, Kristi
Choromanska, Anna
author_facet Dimlioglu, Tolga
Topollai, Kristi
Choromanska, Anna
contents Deep neural network training often exhibits highly anisotropic loss geometry, where a few sharp dominant Hessian directions coexist with a large flatter bulk. Gradients tend to align disproportionately with these dominant directions, although stable progress often requires movement through flatter bulk directions. Estimating the dominant subspace is therefore useful but costly with direct Hessian-based methods. We show that standard Local SGD exposes this geometry through worker disagreement. We theoretically show that the worker-average gap covariance is shaped by stochastic-gradient noise and Hessian curvature, causing workers to disagree along sharp, curvature-sensitive directions. Thus, worker-average gaps provide a cheap Hessian-free estimator of the dominant subspace. Experiments on MLPs, CNNs, and Transformers show that subspaces formed by worker-average gaps capture a substantial fraction of the gradient component lying in the dominant Hessian eigenspace.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Worker Disagreement Reveals Sharp Directions in Local SGD
Dimlioglu, Tolga
Topollai, Kristi
Choromanska, Anna
Machine Learning
Artificial Intelligence
Deep neural network training often exhibits highly anisotropic loss geometry, where a few sharp dominant Hessian directions coexist with a large flatter bulk. Gradients tend to align disproportionately with these dominant directions, although stable progress often requires movement through flatter bulk directions. Estimating the dominant subspace is therefore useful but costly with direct Hessian-based methods. We show that standard Local SGD exposes this geometry through worker disagreement. We theoretically show that the worker-average gap covariance is shaped by stochastic-gradient noise and Hessian curvature, causing workers to disagree along sharp, curvature-sensitive directions. Thus, worker-average gaps provide a cheap Hessian-free estimator of the dominant subspace. Experiments on MLPs, CNNs, and Transformers show that subspaces formed by worker-average gaps capture a substantial fraction of the gradient component lying in the dominant Hessian eigenspace.
title Worker Disagreement Reveals Sharp Directions in Local SGD
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.27739