SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Ce, Wang, Xinghan, Ning, Jiahong, Shi, Yuxuan, Huang, Ning, Yang, Tingting
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908999038795776
author Zheng, Ce
Wang, Xinghan
Ning, Jiahong
Shi, Yuxuan
Huang, Ning
Yang, Tingting
author_facet Zheng, Ce
Wang, Xinghan
Ning, Jiahong
Shi, Yuxuan
Huang, Ning
Yang, Tingting
contents Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent full-model forward passes across workers, severely limiting decoding throughput. Distributed deployment further aggravates this due to a communication bottleneck: each worker must transmit full token probability distributions per draft token, dominating end-to-end latency. To address these challenges, we introduce speculative decoding to enable parallel LLM processing and propose a top-K compressed transmission scheme with two server-side reconstruction strategies. We theoretically analyze the robustness of our method in terms of local reconstruction error, aggregation bias, and acceptance-rate bias, and derive corresponding bounds. Experiments demonstrate that our scheme achieves high generation fidelity while significantly reducing communication overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25777
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
Zheng, Ce
Wang, Xinghan
Ning, Jiahong
Shi, Yuxuan
Huang, Ning
Yang, Tingting
Signal Processing
Distributed, Parallel, and Cluster Computing
Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent full-model forward passes across workers, severely limiting decoding throughput. Distributed deployment further aggravates this due to a communication bottleneck: each worker must transmit full token probability distributions per draft token, dominating end-to-end latency. To address these challenges, we introduce speculative decoding to enable parallel LLM processing and propose a top-K compressed transmission scheme with two server-side reconstruction strategies. We theoretically analyze the robustness of our method in terms of local reconstruction error, aggregation bias, and acceptance-rate bias, and derive corresponding bounds. Experiments demonstrate that our scheme achieves high generation fidelity while significantly reducing communication overhead.
title SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
topic Signal Processing
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2604.25777