FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yijiang, Dey, Emon, Li, Zilinghan, Raghavan, Krishnan, Madduri, Ravi, Kim, Kibaek
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918530902917120
author Li, Yijiang
Dey, Emon
Li, Zilinghan
Raghavan, Krishnan
Madduri, Ravi
Kim, Kibaek
author_facet Li, Yijiang
Dey, Emon
Li, Zilinghan
Raghavan, Krishnan
Madduri, Ravi
Kim, Kibaek
contents Federated learning (FL) across multiple HPC facilities faces stochastic admission delays from batch schedulers that dominate wall-clock time. Synchronous FL suffers from severe stragglers, while asynchronous FL accumulates stale updates when queues spike. We propose FedQueue, a queue-aware FL protocol that incorporates scheduler delays directly into training and aggregation, which (i) predicts per-facility queue delays online to budget local work, (ii) applies cutoff-based admission that buffers late arrivals to bound staleness, and (iii) performs staleness-aware aggregation to stabilize heterogeneous local workloads. We prove the convergence for non-convex objectives at rate $\mathcal{O}(1/\sqrt{R})$ under bounded staleness, and show that the admission controls yield bounded staleness with high probability under queue-prediction error. Real-world cross-facility deployment of FedQueue shows 20.5% improvement over baseline algorithms. Controlled queue simulations demonstrate robust improvement over the baselines; in particular, up to 60% reduction in time to reach a target accuracy level under high queue variance and non-IID partitions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02125
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
Li, Yijiang
Dey, Emon
Li, Zilinghan
Raghavan, Krishnan
Madduri, Ravi
Kim, Kibaek
Distributed, Parallel, and Cluster Computing
Machine Learning
Federated learning (FL) across multiple HPC facilities faces stochastic admission delays from batch schedulers that dominate wall-clock time. Synchronous FL suffers from severe stragglers, while asynchronous FL accumulates stale updates when queues spike. We propose FedQueue, a queue-aware FL protocol that incorporates scheduler delays directly into training and aggregation, which (i) predicts per-facility queue delays online to budget local work, (ii) applies cutoff-based admission that buffers late arrivals to bound staleness, and (iii) performs staleness-aware aggregation to stabilize heterogeneous local workloads. We prove the convergence for non-convex objectives at rate $\mathcal{O}(1/\sqrt{R})$ under bounded staleness, and show that the admission controls yield bounded staleness with high probability under queue-prediction error. Real-world cross-facility deployment of FedQueue shows 20.5% improvement over baseline algorithms. Controlled queue simulations demonstrate robust improvement over the baselines; in particular, up to 60% reduction in time to reach a target accuracy level under high queue variance and non-IID partitions.
title FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2605.02125