MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Zhaoyuan, Zhang, Zeyu, Lan, Tingfeng, Wang, Zirui, Shen, Haiying, Yang, Juncheng, Cheng, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!