Federated Inference for Heterogeneous LLM Communication and Collaboration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zihan, Li, Zeshen, Yang, Howard H., Quek, Tony Q. S., Park, Jihong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Age-Aware Partial Gradient Update Strategy for Federated Learning Over the Air
von: Du, Ruihao, et al.
Veröffentlicht: (2025)
von: Du, Ruihao, et al.
Veröffentlicht: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
Communication-Efficient Federated Learning by Quantized Variance Reduction for Heterogeneous Wireless Edge Networks
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Robust Federated Fine-Tuning in Heterogeneous Networks with Unreliable Connections: An Aggregation View
von: Wang, Yanmeng, et al.
Veröffentlicht: (2025)
von: Wang, Yanmeng, et al.
Veröffentlicht: (2025)
Accelerating Wireless Distributed Learning via Hybrid Split and Federated Learning Optimization
von: Guo, Kun, et al.
Veröffentlicht: (2025)
von: Guo, Kun, et al.
Veröffentlicht: (2025)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
von: Ma, Chengjie, et al.
Veröffentlicht: (2025)
von: Ma, Chengjie, et al.
Veröffentlicht: (2025)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
Integrated user scheduling and beam steering in over-the-air federated learning for mobile IoT
von: Liu, Shengheng, et al.
Veröffentlicht: (2025)
von: Liu, Shengheng, et al.
Veröffentlicht: (2025)
The Role of Federated Learning in a Wireless World with Foundation Models
von: Chen, Zihan, et al.
Veröffentlicht: (2023)
von: Chen, Zihan, et al.
Veröffentlicht: (2023)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
Secure Communication in the Presence of an RIS-Enhanced Eavesdropper in MIMO Networks
von: Zhang, Gaoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gaoyuan, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
von: Kong, Jie, et al.
Veröffentlicht: (2026)
von: Kong, Jie, et al.
Veröffentlicht: (2026)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
HLoRA: Efficient Federated Learning System for LLM Heterogeneous Fine-Tuning
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Data Heterogeneity-Aware Client Selection for Federated Learning in Wireless Networks
von: Yang, Yanbing, et al.
Veröffentlicht: (2025)
von: Yang, Yanbing, et al.
Veröffentlicht: (2025)
Accelerating Privacy-Preserving Federated Learning in Large-Scale LEO Satellite Systems
von: Guo, Binquan, et al.
Veröffentlicht: (2025)
von: Guo, Binquan, et al.
Veröffentlicht: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
Analysis and Optimization of Wireless Multimodal Federated Learning on Modal Heterogeneity
von: Han, Xuefeng, et al.
Veröffentlicht: (2025)
von: Han, Xuefeng, et al.
Veröffentlicht: (2025)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
von: Xie, Zuan, et al.
Veröffentlicht: (2024)
von: Xie, Zuan, et al.
Veröffentlicht: (2024)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Age-Aware Partial Gradient Update Strategy for Federated Learning Over the Air
von: Du, Ruihao, et al.
Veröffentlicht: (2025) -
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025) -
Communication-Efficient Federated Learning by Quantized Variance Reduction for Heterogeneous Wireless Edge Networks
von: Wang, Shuai, et al.
Veröffentlicht: (2025) -
Robust Federated Fine-Tuning in Heterogeneous Networks with Unreliable Connections: An Aggregation View
von: Wang, Yanmeng, et al.
Veröffentlicht: (2025) -
Accelerating Wireless Distributed Learning via Hybrid Split and Federated Learning Optimization
von: Guo, Kun, et al.
Veröffentlicht: (2025)