Controlled disagreement improves generalization in decentralized training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zesen, Johansson, Mikael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914303577161728
author Wang, Zesen
Johansson, Mikael
author_facet Wang, Zesen
Johansson, Mikael
contents Decentralized training is often regarded as inferior to centralized training because the consensus errors between workers are thought to undermine convergence and generalization, even with homogeneous data distributions. This work challenges this view by introducing decentralized SGD with Adaptive Consensus (DSGD-AC), which intentionally preserves non-vanishing consensus errors through a time-dependent scaling mechanism. We prove that these errors are not random noise but systematically align with the dominant Hessian subspace, acting as structured perturbations that guide optimization toward flatter minima. Across image classification and machine translation benchmarks, DSGD-AC consistently surpasses both standard DSGD and centralized SGD in test accuracy and solution flatness. Together, these results establish consensus errors as a useful implicit regularizer and open a new perspective on the design of decentralized learning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02899
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Controlled disagreement improves generalization in decentralized training
Wang, Zesen
Johansson, Mikael
Machine Learning
Distributed, Parallel, and Cluster Computing
Decentralized training is often regarded as inferior to centralized training because the consensus errors between workers are thought to undermine convergence and generalization, even with homogeneous data distributions. This work challenges this view by introducing decentralized SGD with Adaptive Consensus (DSGD-AC), which intentionally preserves non-vanishing consensus errors through a time-dependent scaling mechanism. We prove that these errors are not random noise but systematically align with the dominant Hessian subspace, acting as structured perturbations that guide optimization toward flatter minima. Across image classification and machine translation benchmarks, DSGD-AC consistently surpasses both standard DSGD and centralized SGD in test accuracy and solution flatness. Together, these results establish consensus errors as a useful implicit regularizer and open a new perspective on the design of decentralized learning algorithms.
title Controlled disagreement improves generalization in decentralized training
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.02899