HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yousefijamarani, Zahra, Wang, Xinglu, Wang, Qian, Heisler, Morgan Lindsay, Shabani, Taha, Gholipour, Niloofar, Yassini, Parham, Chang, Hong, Chen, Kan, Zhang, Qiantao, Bai, Xiaolong, Wang, Jiannan, Xiong, Ying, Zhang, Yong, Fan, Zhenan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing
von: Jang, Daesik, et al.
Veröffentlicht: (2026)
von: Jang, Daesik, et al.
Veröffentlicht: (2026)
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
von: Singh, Gursimran, et al.
Veröffentlicht: (2024)
von: Singh, Gursimran, et al.
Veröffentlicht: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
Enhancing Learned Knowledge in LoRA Adapters Through Efficient Contrastive Decoding on Ascend NPUs
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
DL-PIM: Improving Data Locality in Processing-in-Memory Systems
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
von: Rafiei, Davood, et al.
Veröffentlicht: (2025)
von: Rafiei, Davood, et al.
Veröffentlicht: (2025)
Artificial Intelligence for Operations Research: Revolutionizing the Operations Research Process
von: Fan, Zhenan, et al.
Veröffentlicht: (2024)
von: Fan, Zhenan, et al.
Veröffentlicht: (2024)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
von: Lysenstøen, Christian
Veröffentlicht: (2026)
von: Lysenstøen, Christian
Veröffentlicht: (2026)
Fast Heterogeneous Serving: Scalable Mixed-Scale LLM Allocation for SLO-Constrained Inference
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
von: Shi, Ge, et al.
Veröffentlicht: (2025)
von: Shi, Ge, et al.
Veröffentlicht: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
morgan-heisler/DeckBench: v1.0.0 — DECKBench Initial Release (KDD 2026)
von: dsjang2, et al.
Veröffentlicht: (2026)
von: dsjang2, et al.
Veröffentlicht: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
von: Huang, Kaiyu, et al.
Veröffentlicht: (2025)
von: Huang, Kaiyu, et al.
Veröffentlicht: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
SemAug: Semantically Meaningful Image Augmentations for Object Detection Through Language Grounding
von: Heisler, Morgan, et al.
Veröffentlicht: (2022)
von: Heisler, Morgan, et al.
Veröffentlicht: (2022)
Fair and efficient contribution valuation for vertical federated learning
von: Fan, Zhenan, et al.
Veröffentlicht: (2022)
von: Fan, Zhenan, et al.
Veröffentlicht: (2022)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
von: Liu, Banruo, et al.
Veröffentlicht: (2025)
von: Liu, Banruo, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
Validating Interpretability in siRNA Efficacy Prediction: A Perturbation-Based, Dataset-Aware Protocol
von: Khodagholi, Zahra, et al.
Veröffentlicht: (2026)
von: Khodagholi, Zahra, et al.
Veröffentlicht: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
von: Sedlak, Boris, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
Uso de plantas medicinales en el cuidado de la salud: la producción científica de tesis y disertaciones de enfermería brasileña
von: Elisa Vanessa Heisler
Veröffentlicht: (2015)
von: Elisa Vanessa Heisler
Veröffentlicht: (2015)
“Populism and Democracy.” A review of The Age of Discontent: Populism, Extremis and Conspiracy Theories in Contemporary Democracies By MathewRhodes‐Purdy, RachelNavaree and StephenUtych, Cambridge, UK: Cambridge University Press. 2024. $34.99 (pbk); $110.00 (hbk); $110.00 (ebk)
von: Barbara Schmitter Heisler
Veröffentlicht: (2024)
von: Barbara Schmitter Heisler
Veröffentlicht: (2024)
Ähnliche Einträge
-
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
von: Wang, Qian, et al.
Veröffentlicht: (2025) -
DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing
von: Jang, Daesik, et al.
Veröffentlicht: (2026) -
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
von: Singh, Gursimran, et al.
Veröffentlicht: (2024) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025) -
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)