Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Barnes, L. P., Dytso, Alex, Poor, H. V.
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913194533978112
author Barnes, L. P.
Dytso, Alex
Poor, H. V.
author_facet Barnes, L. P.
Dytso, Alex
Poor, H. V.
contents We consider information-theoretic bounds on expected generalization error for statistical learning problems in a networked setting. In this setting, there are $K$ nodes, each with its own independent dataset, and the models from each node have to be aggregated into a final centralized model. We consider both simple averaging of the models as well as more complicated multi-round algorithms. We give upper bounds on the expected generalization error for a variety of problems, such as those with Bregman divergence or Lipschitz continuous losses, that demonstrate an improved dependence of $1/K$ on the number of nodes. These "per node" bounds are in terms of the mutual information between the training dataset and the trained weights at each node, and are therefore useful in describing the generalization properties inherent to having communication or privacy constraints at each node.
format Preprint
id arxiv_https___arxiv_org_abs_2202_02423
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning
Barnes, L. P.
Dytso, Alex
Poor, H. V.
Information Theory
Machine Learning
We consider information-theoretic bounds on expected generalization error for statistical learning problems in a networked setting. In this setting, there are $K$ nodes, each with its own independent dataset, and the models from each node have to be aggregated into a final centralized model. We consider both simple averaging of the models as well as more complicated multi-round algorithms. We give upper bounds on the expected generalization error for a variety of problems, such as those with Bregman divergence or Lipschitz continuous losses, that demonstrate an improved dependence of $1/K$ on the number of nodes. These "per node" bounds are in terms of the mutual information between the training dataset and the trained weights at each node, and are therefore useful in describing the generalization properties inherent to having communication or privacy constraints at each node.
title Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning
topic Information Theory
Machine Learning
url https://arxiv.org/abs/2202.02423