Floe: Federated Specialization for Real-Time LLM-SLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Chunlin, Tam, Kahou, Wu, Yebo, Zhong, Shuaihang, Li, Li, Lane, Nicholas D., Xu, Chengzhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
by: Wu, Yebo, et al.
Published: (2024)
by: Wu, Yebo, et al.
Published: (2024)
Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting
by: Tian, Chunlin, et al.
Published: (2024)
by: Tian, Chunlin, et al.
Published: (2024)
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
by: Wu, Yebo, et al.
Published: (2026)
by: Wu, Yebo, et al.
Published: (2026)
A Survey on Federated Fine-tuning of Large Language Models
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
by: Wu, Yebo, et al.
Published: (2024)
by: Wu, Yebo, et al.
Published: (2024)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
by: Yebo, Wu, et al.
Published: (2024)
by: Yebo, Wu, et al.
Published: (2024)
Ranking-based Client Selection with Imitation Learning for Efficient Federated Learning
by: Tian, Chunlin, et al.
Published: (2024)
by: Tian, Chunlin, et al.
Published: (2024)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
by: He, Yiyuan, et al.
Published: (2024)
by: He, Yiyuan, et al.
Published: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
by: Liao, Junhan, et al.
Published: (2025)
by: Liao, Junhan, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
by: Chen, Zihan, et al.
Published: (2026)
by: Chen, Zihan, et al.
Published: (2026)
Enshrined Proposer Builder Separation in the presence of Maximal Extractable Value
by: Wang, Yitian, et al.
Published: (2026)
by: Wang, Yitian, et al.
Published: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
by: Lin, Yanying, et al.
Published: (2025)
by: Lin, Yanying, et al.
Published: (2025)
Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
Decentralized Semantic Federated Learning for Real-Time Public Safety Tasks: Challenges, Methods, and Directions
by: Li, Baosheng, et al.
Published: (2025)
by: Li, Baosheng, et al.
Published: (2025)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
by: Zhang, Songge, et al.
Published: (2026)
by: Zhang, Songge, et al.
Published: (2026)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
by: He, Yiyuan, et al.
Published: (2025)
by: He, Yiyuan, et al.
Published: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
by: Chen, Xiaopei, et al.
Published: (2025)
by: Chen, Xiaopei, et al.
Published: (2025)
DynaShard: Secure and Adaptive Blockchain Sharding Protocol with Hybrid Consensus and Dynamic Shard Management
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
by: Zhan, Huiyou, et al.
Published: (2025)
by: Zhan, Huiyou, et al.
Published: (2025)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
by: Bai, Haoyu, et al.
Published: (2024)
by: Bai, Haoyu, et al.
Published: (2024)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
by: Babaei, Amir Fakhim, et al.
Published: (2025)
by: Babaei, Amir Fakhim, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
Similar Items
-
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
by: Wu, Yebo, et al.
Published: (2024) -
Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting
by: Tian, Chunlin, et al.
Published: (2024) -
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
by: Wu, Yebo, et al.
Published: (2026) -
A Survey on Federated Fine-tuning of Large Language Models
by: Wu, Yebo, et al.
Published: (2025) -
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
by: Wu, Yebo, et al.
Published: (2024)