SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Ce, Wang, Xinghan, Ning, Jiahong, Shi, Yuxuan, Huang, Ning, Yang, Tingting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
by: Zheng, Ce, et al.
Published: (2026)
by: Zheng, Ce, et al.
Published: (2026)
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
IRS Aided Federated Learning: Multiple Access and Fundamental Tradeoff
by: Chen, Guangji, et al.
Published: (2024)
by: Chen, Guangji, et al.
Published: (2024)
Over-the-Air Federated Learning with Phase Noise: Analysis and Countermeasures
by: Dahl, Martin, et al.
Published: (2024)
by: Dahl, Martin, et al.
Published: (2024)
Over-the-air Federated Policy Gradient
by: Yang, Huiwen, et al.
Published: (2023)
by: Yang, Huiwen, et al.
Published: (2023)
PowerTrain: Fast, Generalizable Time and Power Prediction Models to Optimize DNN Training on Accelerated Edges
by: K., Prashanthi S., et al.
Published: (2024)
by: K., Prashanthi S., et al.
Published: (2024)
Compressed Bayesian Federated Learning for Reliable Passive Radio Sensing in Industrial IoT
by: Barbieri, Luca, et al.
Published: (2024)
by: Barbieri, Luca, et al.
Published: (2024)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
by: Qin, Shengling, et al.
Published: (2025)
by: Qin, Shengling, et al.
Published: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
Design of A Low-Latency and Parallelizable SVD Dataflow Architecture on FPGA
by: Du, Fangqiang, et al.
Published: (2025)
by: Du, Fangqiang, et al.
Published: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
by: Ning, Jiahong, et al.
Published: (2025)
by: Ning, Jiahong, et al.
Published: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025)
by: Xu, Jiaming, et al.
Published: (2025)
An Experimental Exploration of In-Memory Computing for Multi-Layer Perceptrons
by: Carrinho, Pedro, et al.
Published: (2025)
by: Carrinho, Pedro, et al.
Published: (2025)
RASC: Region-Aware Self-Calibration for Dense 2D Sensor Arrays
by: Ma, Yinglei, et al.
Published: (2026)
by: Ma, Yinglei, et al.
Published: (2026)
Computation Offloading Strategies in Integrated Terrestrial and Non-Terrestrial Networks
by: Mohsin, Muhammad Ahmed, et al.
Published: (2025)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2025)
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis
by: Zhang, Xueyao, et al.
Published: (2024)
by: Zhang, Xueyao, et al.
Published: (2024)
Parallel-in-Time Kalman Smoothing Using Orthogonal Transformations
by: Gargir, Shahaf, et al.
Published: (2025)
by: Gargir, Shahaf, et al.
Published: (2025)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
by: Lv, Cunchi, et al.
Published: (2025)
by: Lv, Cunchi, et al.
Published: (2025)
Convergence of Agnostic Federated Averaging
by: Herlock, et al.
Published: (2025)
by: Herlock, et al.
Published: (2025)
SignSGD with Federated Voting
by: Park, Chanho, et al.
Published: (2024)
by: Park, Chanho, et al.
Published: (2024)
Energy-Aware Federated Learning in Satellite Constellations
by: Razmi, Nasrin, et al.
Published: (2024)
by: Razmi, Nasrin, et al.
Published: (2024)
Sparse Incremental Aggregation in Multi-Hop Federated Learning
by: Mukherjee, Sourav, et al.
Published: (2024)
by: Mukherjee, Sourav, et al.
Published: (2024)
Noise-Robust and Resource-Efficient ADMM-based Federated Learning
by: Lari, Ehsan, et al.
Published: (2024)
by: Lari, Ehsan, et al.
Published: (2024)
Real-Time Diagnostic Integrity Meets Efficiency: A Novel Platform-Agnostic Architecture for Physiological Signal Compression
by: Vora, Neel R, et al.
Published: (2023)
by: Vora, Neel R, et al.
Published: (2023)
Communication Efficient ConFederated Learning: An Event-Triggered SAGA Approach
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Bayes-Split-Edge: Bayesian Optimization for Constrained Collaborative Inference in Wireless Edge Systems
by: Safaeipour, Fatemeh Zahra, et al.
Published: (2025)
by: Safaeipour, Fatemeh Zahra, et al.
Published: (2025)
A Novel Framework of Horizontal-Vertical Hybrid Federated Learning for EdgeIoT
by: Li, Kai, et al.
Published: (2024)
by: Li, Kai, et al.
Published: (2024)
Meta-Federated Learning: A Novel Approach for Real-Time Traffic Flow Management
by: Johnson, Bob, et al.
Published: (2025)
by: Johnson, Bob, et al.
Published: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
Temporal Predictive Coding for Gradient Compression in Distributed Learning
by: Edin, Adrian, et al.
Published: (2024)
by: Edin, Adrian, et al.
Published: (2024)
FedAQ: Communication-Efficient Federated Edge Learning via Joint Uplink and Downlink Adaptive Quantization
by: Qu, Linping, et al.
Published: (2024)
by: Qu, Linping, et al.
Published: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
A Distributed Plug-and-Play MCMC Algorithm for High-Dimensional Inverse Problems
by: Bouton, Maxime, et al.
Published: (2025)
by: Bouton, Maxime, et al.
Published: (2025)
Similar Items
-
Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
by: Zheng, Ce, et al.
Published: (2026) -
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
by: Zhang, Kai, et al.
Published: (2025) -
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025) -
IRS Aided Federated Learning: Multiple Access and Fundamental Tradeoff
by: Chen, Guangji, et al.
Published: (2024) -
Over-the-Air Federated Learning with Phase Noise: Analysis and Countermeasures
by: Dahl, Martin, et al.
Published: (2024)