Gespeichert in:
| Hauptverfasser: | Yu, Lingfan, Lin, Jinkun, Li, Jinyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2312.05516 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
von: Zhao, Adrian, et al.
Veröffentlicht: (2026)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
Understanding Stragglers in Large Model Training Using What-if Analysis
von: Lin, Jinkun, et al.
Veröffentlicht: (2025)
von: Lin, Jinkun, et al.
Veröffentlicht: (2025)
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
von: Shen, Ao, et al.
Veröffentlicht: (2024)
von: Shen, Ao, et al.
Veröffentlicht: (2024)
Towards Resiliency in Large Language Model Serving with KevlarFlow
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
von: Chen, Zixi, et al.
Veröffentlicht: (2026)
von: Chen, Zixi, et al.
Veröffentlicht: (2026)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
von: Du, Zhixu, et al.
Veröffentlicht: (2023)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
MUSE: Multi-Tenant Model Serving With Seamless Model Updates
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
EdgeServe: A Streaming System for Decentralized Model Serving
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
von: Mei, Yixuan, et al.
Veröffentlicht: (2024)
von: Mei, Yixuan, et al.
Veröffentlicht: (2024)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
von: Ma, Jeff J., et al.
Veröffentlicht: (2025)
von: Ma, Jeff J., et al.
Veröffentlicht: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Serving Large Language Models on Huawei CloudMatrix384
von: Zuo, Pengfei, et al.
Veröffentlicht: (2025)
von: Zuo, Pengfei, et al.
Veröffentlicht: (2025)
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2024)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2024)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
von: Yu, Jiahuan, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
von: Chen, Shaoyuan, et al.
Veröffentlicht: (2024)
von: Chen, Shaoyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
von: Zhao, Adrian, et al.
Veröffentlicht: (2026) -
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024) -
Understanding Stragglers in Large Model Training Using What-if Analysis
von: Lin, Jinkun, et al.
Veröffentlicht: (2025) -
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024) -
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)