FedSQ: Optimized Weight Averaging via Fixed Gating

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pérez-Corral, Cristian, Mestre, Jose I., Fernández-Hernández, Alberto, Dolz, Manuel F., Duato, José, Quintana-Ortí, Enrique S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917382877872128
author Pérez-Corral, Cristian
Mestre, Jose I.
Fernández-Hernández, Alberto
Dolz, Manuel F.
Duato, José
Quintana-Ortí, Enrique S.
author_facet Pérez-Corral, Cristian
Mestre, Jose I.
Fernández-Hernández, Alberto
Dolz, Manuel F.
Duato, José
Quintana-Ortí, Enrique S.
contents Federated learning (FL) enables collaborative training across organizations without sharing raw data, but it is hindered by statistical heterogeneity (non-i.i.d.\ client data) and by instability of naive weight averaging under client drift. In many cross-silo deployments, FL is warm-started from a strong pretrained backbone (e.g., ImageNet-1K) and then adapted to local domains. Motivated by recent evidence that ReLU-like gating regimes (structural knowledge) stabilize earlier than the remaining parameter values (quantitative knowledge), we propose FedSQ (Federated Structural-Quantitative learning), a transfer-initialized neural federated procedure based on a DualCopy, piecewise-linear view of deep networks. FedSQ freezes a structural copy of the pretrained model to induce fixed binary gating masks during federated fine-tuning, while only a quantitative copy is optimized locally and aggregated across rounds. Fixing the gating reduces learning to within-regime affine refinements, which stabilizes aggregation under heterogeneous partitions. Experiments on two convolutional neural network backbones under i.i.d.\ and Dirichlet splits show that FedSQ improves robustness and can reduce rounds-to-best validation performance relative to standard baselines while preserving accuracy in the transfer setting.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02990
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FedSQ: Optimized Weight Averaging via Fixed Gating
Pérez-Corral, Cristian
Mestre, Jose I.
Fernández-Hernández, Alberto
Dolz, Manuel F.
Duato, José
Quintana-Ortí, Enrique S.
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Federated learning (FL) enables collaborative training across organizations without sharing raw data, but it is hindered by statistical heterogeneity (non-i.i.d.\ client data) and by instability of naive weight averaging under client drift. In many cross-silo deployments, FL is warm-started from a strong pretrained backbone (e.g., ImageNet-1K) and then adapted to local domains. Motivated by recent evidence that ReLU-like gating regimes (structural knowledge) stabilize earlier than the remaining parameter values (quantitative knowledge), we propose FedSQ (Federated Structural-Quantitative learning), a transfer-initialized neural federated procedure based on a DualCopy, piecewise-linear view of deep networks. FedSQ freezes a structural copy of the pretrained model to induce fixed binary gating masks during federated fine-tuning, while only a quantitative copy is optimized locally and aggregated across rounds. Fixing the gating reduces learning to within-regime affine refinements, which stabilizes aggregation under heterogeneous partitions. Experiments on two convolutional neural network backbones under i.i.d.\ and Dirichlet splits show that FedSQ improves robustness and can reduce rounds-to-best validation performance relative to standard baselines while preserving accuracy in the transfer setting.
title FedSQ: Optimized Weight Averaging via Fixed Gating
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2604.02990