Hierarchical Autoscaling for Large Language Model Serving with Chiron
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patke, Archit, Reddy, Dhemath, Jha, Saurabh, Narayanaswami, Chandra, Kalbarczyk, Zbigniew, Iyer, Ravishankar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Queue management for slo-oriented large language model serving
von: Patke, Archit, et al.
Veröffentlicht: (2024)
von: Patke, Archit, et al.
Veröffentlicht: (2024)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
Mutiny! How does Kubernetes fail, and what can we do about it?
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
PRAXIS: Integrating Program Analysis with Observability for Root-Cause Analysis
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
von: Fang, Fei, et al.
Veröffentlicht: (2025)
von: Fang, Fei, et al.
Veröffentlicht: (2025)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
Multi-Objective Optimization of Consumer Group Autoscaling in Message Broker Systems
von: Landau, Diogo, et al.
Veröffentlicht: (2024)
von: Landau, Diogo, et al.
Veröffentlicht: (2024)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Scaling Performance of Large Language Model Pretraining
von: Interrante-Grant, Alexander, et al.
Veröffentlicht: (2025)
von: Interrante-Grant, Alexander, et al.
Veröffentlicht: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
von: Yang, Lingyun, et al.
Veröffentlicht: (2026)
von: Yang, Lingyun, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
Can Large Language Models Write Parallel Code?
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
Efficient Multi-Model Orchestration for Self-Hosted Large Language Models
von: Vangala, Bhanu Prakash, et al.
Veröffentlicht: (2025)
von: Vangala, Bhanu Prakash, et al.
Veröffentlicht: (2025)
HPC-Coder: Modeling Parallel Programs using Large Language Models
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Queue management for slo-oriented large language model serving
von: Patke, Archit, et al.
Veröffentlicht: (2024) -
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025) -
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
von: Cui, Shengkun, et al.
Veröffentlicht: (2025) -
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
von: Qiu, Haoran, et al.
Veröffentlicht: (2024) -
Mutiny! How does Kubernetes fail, and what can we do about it?
von: Barletta, Marco, et al.
Veröffentlicht: (2024)