ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Ge, Sadri, Hanieh, Wang, Qian, Zhang, Yu, Xiong, Ying, Zhang, Yong, Fan, Zhenan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
von: Singh, Gursimran, et al.
Veröffentlicht: (2025)
von: Singh, Gursimran, et al.
Veröffentlicht: (2025)
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
von: Yu, Timothy Tin Long, et al.
Veröffentlicht: (2026)
von: Yu, Timothy Tin Long, et al.
Veröffentlicht: (2026)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
HLoRA: Efficient Federated Learning System for LLM Heterogeneous Fine-Tuning
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
EcoLoRA: Communication-Efficient Federated Fine-Tuning of Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2025)
von: Liu, Han, et al.
Veröffentlicht: (2025)
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
von: Ding, Chuntao, et al.
Veröffentlicht: (2024)
von: Ding, Chuntao, et al.
Veröffentlicht: (2024)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
von: Yi, Mengjun, et al.
Veröffentlicht: (2025)
von: Yi, Mengjun, et al.
Veröffentlicht: (2025)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
von: Wang, Minghe, et al.
Veröffentlicht: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
von: Singh, Gursimran, et al.
Veröffentlicht: (2025) -
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
von: Yu, Timothy Tin Long, et al.
Veröffentlicht: (2026) -
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025) -
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)