CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
Fuente:
arXiv
Salvato in:
| Autori principali: | Suo, Jiashun, Liao, Xiaojian, Xiao, Limin, Ruan, Li, Wang, Jinquan, Su, Xiao, Huo, Zhisheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
di: Wang, Jinquan, et al.
Pubblicazione: (2025)
di: Wang, Jinquan, et al.
Pubblicazione: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
di: Ng, Nathan, et al.
Pubblicazione: (2026)
di: Ng, Nathan, et al.
Pubblicazione: (2026)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
di: Sun, Tingyang, et al.
Pubblicazione: (2026)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)
di: Shen, Zixu, et al.
Pubblicazione: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?
di: Dou, Xinglei, et al.
Pubblicazione: (2025)
di: Dou, Xinglei, et al.
Pubblicazione: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
di: Scheinert, Dominik, et al.
Pubblicazione: (2023)
di: Scheinert, Dominik, et al.
Pubblicazione: (2023)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining
di: Singapuram, Sanjay Sri Vallabh, et al.
Pubblicazione: (2025)
di: Singapuram, Sanjay Sri Vallabh, et al.
Pubblicazione: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
di: Xiao, Bin, et al.
Pubblicazione: (2024)
di: Xiao, Bin, et al.
Pubblicazione: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
di: Arif, Moiz, et al.
Pubblicazione: (2026)
di: Arif, Moiz, et al.
Pubblicazione: (2026)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024)
di: He, Jiaao, et al.
Pubblicazione: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
CoNST: Code Generator for Sparse Tensor Networks
di: Raje, Saurabh, et al.
Pubblicazione: (2024)
di: Raje, Saurabh, et al.
Pubblicazione: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
di: Shan, Baodi, et al.
Pubblicazione: (2024)
di: Shan, Baodi, et al.
Pubblicazione: (2024)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2025)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
di: Qararyah, Fareed, et al.
Pubblicazione: (2024)
di: Qararyah, Fareed, et al.
Pubblicazione: (2024)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
di: Mathews, Dhanya R, et al.
Pubblicazione: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
di: Muhammad, Said, et al.
Pubblicazione: (2025)
di: Muhammad, Said, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
di: Wang, Jinquan, et al.
Pubblicazione: (2025) -
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
di: Ng, Nathan, et al.
Pubblicazione: (2026) -
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025) -
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
di: Sun, Tingyang, et al.
Pubblicazione: (2026) -
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)