MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Chunyu, Pan, Yi, Cui, Weihao, Chen, Quan, Zhang, Shulai, He, Bingsheng, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
von: Xu, Ao, et al.
Veröffentlicht: (2025)
von: Xu, Ao, et al.
Veröffentlicht: (2025)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
A Reinforcement Learning-Driven Task Scheduling Algorithm for Multi-Tenant Distributed Systems
von: Zhang, Xiaopei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaopei, et al.
Veröffentlicht: (2025)
Caching Aided Multi-Tenant Serverless Computing
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
Guardian: Safe GPU Sharing in Multi-Tenant Environments
von: Pavlidakis, Manos, et al.
Veröffentlicht: (2024)
von: Pavlidakis, Manos, et al.
Veröffentlicht: (2024)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
Trabant: A Serverless Architecture for Multi-Tenant Orbital Edge Computing
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2025)
von: Pfandzelter, Tobias, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
HLoRA: Efficient Federated Learning System for LLM Heterogeneous Fine-Tuning
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Multi Armed Bandit Algorithms Based Virtual Machine Allocation Policy for Security in Multi-Tenant Distributed Systems
von: Patil, Pravin, et al.
Veröffentlicht: (2024)
von: Patil, Pravin, et al.
Veröffentlicht: (2024)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
von: Griggs, Tyler, et al.
Veröffentlicht: (2026)
von: Griggs, Tyler, et al.
Veröffentlicht: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning
von: Lyu, Chenghao, et al.
Veröffentlicht: (2024)
von: Lyu, Chenghao, et al.
Veröffentlicht: (2024)
Split Fine-Tuning for Large Language Models in Wireless Networks
von: Zhang, Songge, et al.
Veröffentlicht: (2025)
von: Zhang, Songge, et al.
Veröffentlicht: (2025)
Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation
von: Zhang, Zikai, et al.
Veröffentlicht: (2026)
von: Zhang, Zikai, et al.
Veröffentlicht: (2026)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
von: Moore, H., et al.
Veröffentlicht: (2026)
von: Moore, H., et al.
Veröffentlicht: (2026)
Datacenter Energy Optimized Power Profiles
von: Narayanaswamy, Sreedhar, et al.
Veröffentlicht: (2025)
von: Narayanaswamy, Sreedhar, et al.
Veröffentlicht: (2025)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Symbiosis: Multi-Adapter Inference and Fine-Tuning
von: Gupta, Saransh, et al.
Veröffentlicht: (2025)
von: Gupta, Saransh, et al.
Veröffentlicht: (2025)
Equilibria: Fair Multi-Tenant CXL Memory Tiering At Scale
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
von: Yu, Timothy Tin Long, et al.
Veröffentlicht: (2026)
von: Yu, Timothy Tin Long, et al.
Veröffentlicht: (2026)
MUSE: Multi-Tenant Model Serving With Seamless Model Updates
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024) -
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025) -
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024) -
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
von: Xu, Ao, et al.
Veröffentlicht: (2025)