FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
Fuente:
arXiv
Guardado en:
| Autores principales: | Oliaro, Gabriele, Miao, Xupeng, Cheng, Xinhao, Kada, Vineeth, Wu, Mengdi, Gao, Ruohan, Huang, Yingyi, Delacourt, Remi, Yang, April, Wang, Yingcheng, Unger, Colin, Jia, Zhihao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025)
por: Li, Zikun, et al.
Publicado: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
por: Miao, Xupeng, et al.
Publicado: (2023)
por: Miao, Xupeng, et al.
Publicado: (2023)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024)
por: Shen, Haiying, et al.
Publicado: (2024)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
por: Chen, Bocheng, et al.
Publicado: (2024)
por: Chen, Bocheng, et al.
Publicado: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
por: Shen, Haiying, et al.
Publicado: (2025)
por: Shen, Haiying, et al.
Publicado: (2025)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
por: Mei, Yixuan, et al.
Publicado: (2026)
por: Mei, Yixuan, et al.
Publicado: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
por: Nie, Chengyi, et al.
Publicado: (2024)
por: Nie, Chengyi, et al.
Publicado: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
por: Miao, Xupeng, et al.
Publicado: (2023)
por: Miao, Xupeng, et al.
Publicado: (2023)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
por: Zhang, Jiahao, et al.
Publicado: (2026)
por: Zhang, Jiahao, et al.
Publicado: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
por: Lysenstøen, Christian
Publicado: (2026)
por: Lysenstøen, Christian
Publicado: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
por: Zhang, Wei, et al.
Publicado: (2025)
por: Zhang, Wei, et al.
Publicado: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
por: Huang, Shaoyuan, et al.
Publicado: (2026)
por: Huang, Shaoyuan, et al.
Publicado: (2026)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
por: Li, Yufei, et al.
Publicado: (2025)
por: Li, Yufei, et al.
Publicado: (2025)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
por: Xu, Yechen, et al.
Publicado: (2024)
por: Xu, Yechen, et al.
Publicado: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
por: Zhu, Kan, et al.
Publicado: (2025)
por: Zhu, Kan, et al.
Publicado: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
por: Chen, Siyuan, et al.
Publicado: (2025)
por: Chen, Siyuan, et al.
Publicado: (2025)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
por: Du, Yin, et al.
Publicado: (2026)
por: Du, Yin, et al.
Publicado: (2026)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
por: Xu, Ao, et al.
Publicado: (2025)
por: Xu, Ao, et al.
Publicado: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
por: Zhang, Biao, et al.
Publicado: (2024)
por: Zhang, Biao, et al.
Publicado: (2024)
Past-Future Scheduler for LLM Serving under SLA Guarantees
por: Gong, Ruihao, et al.
Publicado: (2025)
por: Gong, Ruihao, et al.
Publicado: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
por: Lin, Yanying, et al.
Publicado: (2025)
por: Lin, Yanying, et al.
Publicado: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
por: Lai, Ruiqi, et al.
Publicado: (2025)
por: Lai, Ruiqi, et al.
Publicado: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
por: Cheng, Ke, et al.
Publicado: (2024)
por: Cheng, Ke, et al.
Publicado: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
por: Mei, Yixuan, et al.
Publicado: (2024)
por: Mei, Yixuan, et al.
Publicado: (2024)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
por: Guo, Tianyu, et al.
Publicado: (2025)
por: Guo, Tianyu, et al.
Publicado: (2025)
Where does an LLM begin computing an instruction?
por: Pola, Aditya, et al.
Publicado: (2025)
por: Pola, Aditya, et al.
Publicado: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving
por: Liu, Xunzhuo, et al.
Publicado: (2026)
por: Liu, Xunzhuo, et al.
Publicado: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
por: Wang, Qipeng
Publicado: (2026)
por: Wang, Qipeng
Publicado: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
por: Chow, Will
Publicado: (2025)
por: Chow, Will
Publicado: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
por: Ma, Chenxiang, et al.
Publicado: (2025)
por: Ma, Chenxiang, et al.
Publicado: (2025)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
por: Yang, Hao, et al.
Publicado: (2026)
por: Yang, Hao, et al.
Publicado: (2026)
Learning Dynamics of LLM Finetuning
por: Ren, Yi, et al.
Publicado: (2024)
por: Ren, Yi, et al.
Publicado: (2024)
AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
por: Huang, Kaiyu, et al.
Publicado: (2025)
por: Huang, Kaiyu, et al.
Publicado: (2025)
Ejemplares similares
-
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025) -
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
por: Miao, Xupeng, et al.
Publicado: (2023) -
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024) -
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
por: Chen, Bocheng, et al.
Publicado: (2024) -
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)