Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Tiancheng, Wang, Chenxi, Cao, Ting, Qin, Jin, Chen, Lei, Xiao, Xinyu, Hu, Junhao, Tian, Hongliang, Yan, Shoumeng, Cui, Huimin, Chen, Quan, Xie, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
di: Han, Mingcong, et al.
Pubblicazione: (2024)
di: Han, Mingcong, et al.
Pubblicazione: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
di: Aguilera, Marcos K., et al.
Pubblicazione: (2022)
di: Aguilera, Marcos K., et al.
Pubblicazione: (2022)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
di: Huang, Shaoyuan, et al.
Pubblicazione: (2026)
di: Huang, Shaoyuan, et al.
Pubblicazione: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
di: Zhang, Mingjun, et al.
Pubblicazione: (2025)
di: Zhang, Mingjun, et al.
Pubblicazione: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
di: Xu, Ao, et al.
Pubblicazione: (2025)
di: Xu, Ao, et al.
Pubblicazione: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
di: Liu, Yi, et al.
Pubblicazione: (2025)
di: Liu, Yi, et al.
Pubblicazione: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
di: Zhao, Han, et al.
Pubblicazione: (2024)
di: Zhao, Han, et al.
Pubblicazione: (2024)
Improved Methods of Task Assignment and Resource Allocation with Preemption in Edge Computing Systems
di: Rublein, Caroline, et al.
Pubblicazione: (2024)
di: Rublein, Caroline, et al.
Pubblicazione: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
di: Sedlak, Boris, et al.
Pubblicazione: (2024)
di: Sedlak, Boris, et al.
Pubblicazione: (2024)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
di: Jeon, Beomyeol, et al.
Pubblicazione: (2024)
di: Jeon, Beomyeol, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
di: Chow, Will
Pubblicazione: (2025)
di: Chow, Will
Pubblicazione: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
di: Xuan, Mo, et al.
Pubblicazione: (2025)
di: Xuan, Mo, et al.
Pubblicazione: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
di: Gao, Wei, et al.
Pubblicazione: (2026)
di: Gao, Wei, et al.
Pubblicazione: (2026)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
di: Nozaki, Ai, et al.
Pubblicazione: (2026)
di: Nozaki, Ai, et al.
Pubblicazione: (2026)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
di: Peng, Haosong, et al.
Pubblicazione: (2024)
di: Peng, Haosong, et al.
Pubblicazione: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
di: Chen, Siyuan, et al.
Pubblicazione: (2025)
di: Chen, Siyuan, et al.
Pubblicazione: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
di: Zhao, Chen, et al.
Pubblicazione: (2024)
di: Zhao, Chen, et al.
Pubblicazione: (2024)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
di: Sun, Desen, et al.
Pubblicazione: (2025)
di: Sun, Desen, et al.
Pubblicazione: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
Heat: Satellite's meat is GPU's poison
di: Yuan, Zhehu, et al.
Pubblicazione: (2024)
di: Yuan, Zhehu, et al.
Pubblicazione: (2024)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
di: Wu, Shixun, et al.
Pubblicazione: (2025)
di: Wu, Shixun, et al.
Pubblicazione: (2025)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
di: Xu, Zhihao, et al.
Pubblicazione: (2025)
di: Xu, Zhihao, et al.
Pubblicazione: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
di: Zhang, WenZheng, et al.
Pubblicazione: (2024)
di: Zhang, WenZheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
di: Duan, Jiaang, et al.
Pubblicazione: (2025) -
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
di: Cheng, Ke, et al.
Pubblicazione: (2024) -
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
di: Gu, Jianfeng, et al.
Pubblicazione: (2025) -
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
di: Hu, Tiancheng, et al.
Pubblicazione: (2026) -
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
di: Han, Mingcong, et al.
Pubblicazione: (2024)