CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Suyi, Lu, Hanfeng, Wu, Tianyuan, Yu, Minchen, Weng, Qizhen, Chen, Xusheng, Shan, Yizhou, Yuan, Binhang, Wang, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
por: Chen, Hongyu, et al.
Publicado: (2026)
por: Chen, Hongyu, et al.
Publicado: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
por: Hu, Cunchen, et al.
Publicado: (2024)
por: Hu, Cunchen, et al.
Publicado: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
por: Wang, Weiye, et al.
Publicado: (2026)
por: Wang, Weiye, et al.
Publicado: (2026)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
por: Peng, You, et al.
Publicado: (2026)
por: Peng, You, et al.
Publicado: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
por: Hu, Junhao, et al.
Publicado: (2025)
por: Hu, Junhao, et al.
Publicado: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
por: Ni, Yinan, et al.
Publicado: (2025)
por: Ni, Yinan, et al.
Publicado: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
por: Wang, Shao, et al.
Publicado: (2026)
por: Wang, Shao, et al.
Publicado: (2026)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
por: Jiang, Youhe, et al.
Publicado: (2026)
por: Jiang, Youhe, et al.
Publicado: (2026)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
por: Li, Allison, et al.
Publicado: (2025)
por: Li, Allison, et al.
Publicado: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
por: Cao, Jiahe, et al.
Publicado: (2026)
por: Cao, Jiahe, et al.
Publicado: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
por: Li, Yueying, et al.
Publicado: (2025)
por: Li, Yueying, et al.
Publicado: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
por: Nguyen, Thanh-Tung, et al.
Publicado: (2025)
por: Nguyen, Thanh-Tung, et al.
Publicado: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
por: He, Yiyuan, et al.
Publicado: (2024)
por: He, Yiyuan, et al.
Publicado: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
por: He, Wenhao, et al.
Publicado: (2026)
por: He, Wenhao, et al.
Publicado: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
por: Liu, Yifei, et al.
Publicado: (2025)
por: Liu, Yifei, et al.
Publicado: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
por: Duan, Jiangfei, et al.
Publicado: (2024)
por: Duan, Jiangfei, et al.
Publicado: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
por: Nie, Chengyi, et al.
Publicado: (2024)
por: Nie, Chengyi, et al.
Publicado: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
por: Sui, Yifan, et al.
Publicado: (2025)
por: Sui, Yifan, et al.
Publicado: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024)
por: Shen, Haiying, et al.
Publicado: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
por: Liu, Jianmin, et al.
Publicado: (2025)
por: Liu, Jianmin, et al.
Publicado: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
por: Huang, Heyang, et al.
Publicado: (2025)
por: Huang, Heyang, et al.
Publicado: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
por: Wolfrath, Joel, et al.
Publicado: (2025)
por: Wolfrath, Joel, et al.
Publicado: (2025)
AutoRank: MCDA Based Rank Personalization for LoRA-Enabled Distributed Learning
por: Chen, Shuaijun, et al.
Publicado: (2024)
por: Chen, Shuaijun, et al.
Publicado: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
por: Zhou, Qihui, et al.
Publicado: (2025)
por: Zhou, Qihui, et al.
Publicado: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
por: Zheng, Wanyi, et al.
Publicado: (2025)
por: Zheng, Wanyi, et al.
Publicado: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
por: Gao, Luyao, et al.
Publicado: (2025)
por: Gao, Luyao, et al.
Publicado: (2025)
Ejemplares similares
-
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026) -
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
por: Chen, Hongyu, et al.
Publicado: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023) -
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
por: Hu, Cunchen, et al.
Publicado: (2024)