Queue management for slo-oriented large language model serving
Fuente:
arXiv
Saved in:
| Main Authors: | Patke, Archit, Reddy, Dhemath, Jha, Saurabh, Qiu, Haoran, Pinto, Christian, Narayanaswami, Chandra, Kalbarczyk, Zbigniew, Iyer, Ravishankar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
by: Qiu, Haoran, et al.
Published: (2024)
by: Qiu, Haoran, et al.
Published: (2024)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
Mutiny! How does Kubernetes fail, and what can we do about it?
by: Barletta, Marco, et al.
Published: (2024)
by: Barletta, Marco, et al.
Published: (2024)
PRAXIS: Integrating Program Analysis with Observability for Root-Cause Analysis
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025)
by: Shetty, Chirag, et al.
Published: (2025)
QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026)
by: Laju, Marco, et al.
Published: (2026)
Highly-Efficient Persistent FIFO Queues
by: Fatourou, Panagiota, et al.
Published: (2024)
by: Fatourou, Panagiota, et al.
Published: (2024)
Aggregating Funnels for Faster Fetch&Add and Queues
by: Roh, Younghun, et al.
Published: (2024)
by: Roh, Younghun, et al.
Published: (2024)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
by: Giannoula, Christina, et al.
Published: (2024)
by: Giannoula, Christina, et al.
Published: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Optimizing FaaS Platforms for MCP-enabled Agentic Workflows
by: Kulkarni, Varad, et al.
Published: (2026)
by: Kulkarni, Varad, et al.
Published: (2026)
Quantum-Enhanced Distributed Sensor Fusion: Lower Bounds on Aggregation from Projection Noise to Heisenberg-Limited Byzantine-Tolerant Networks
by: Iyer, Vasanth, et al.
Published: (2026)
by: Iyer, Vasanth, et al.
Published: (2026)
Transforming Lock-free Linked Lists into Distributed Lock-free Linked Lists
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
by: Ravishankar, Raaghav, et al.
Published: (2024)
by: Ravishankar, Raaghav, et al.
Published: (2024)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
FedQueue: Queue-Aware Federated Learning for Cross-Facility HPC Training
by: Li, Yijiang, et al.
Published: (2026)
by: Li, Yijiang, et al.
Published: (2026)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
by: Chen, Yinfang, et al.
Published: (2025)
by: Chen, Yinfang, et al.
Published: (2025)
Stateless Snowflake: A Cloud-Agnostic Distributed ID Generator Using Network-Derived Identity
by: Chinthareddy, Manideep Reddy
Published: (2025)
by: Chinthareddy, Manideep Reddy
Published: (2025)
Quantum Mini-Apps: A Framework for Developing and Benchmarking Quantum-HPC Applications
by: Saurabh, Nishant, et al.
Published: (2024)
by: Saurabh, Nishant, et al.
Published: (2024)
MQFQ-Sticky: Fair Queueing For Serverless GPU Functions
by: Fuerst, Alexander, et al.
Published: (2025)
by: Fuerst, Alexander, et al.
Published: (2025)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
by: Dai, Yinwei, et al.
Published: (2025)
by: Dai, Yinwei, et al.
Published: (2025)
Engineering MultiQueues: Fast Relaxed Concurrent Priority Queues
by: Williams, Marvin, et al.
Published: (2025)
by: Williams, Marvin, et al.
Published: (2025)
OMP-Engineer: Bridging Syntax Analysis and In-Context Learning for Efficient Automated OpenMP Parallelization
by: Wang, Weidong, et al.
Published: (2024)
by: Wang, Weidong, et al.
Published: (2024)
Memory Bounds for Concurrent Bounded Queues
by: Aksenov, Vitaly, et al.
Published: (2021)
by: Aksenov, Vitaly, et al.
Published: (2021)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Retrofitting Service Dependency Discovery in Distributed Systems
by: Landau, Diogo, et al.
Published: (2025)
by: Landau, Diogo, et al.
Published: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
by: Banchelli, Fabio, et al.
Published: (2025)
by: Banchelli, Fabio, et al.
Published: (2025)
DEVS/SOA: A Cross-Platform Framework for Net-centric Modeling and Simulation in DEVS Unified Process
by: Mittal, Saurabh, et al.
Published: (2024)
by: Mittal, Saurabh, et al.
Published: (2024)
Optimal Checkpoint Interval with Availability as an Objective Function
by: Saxena, Nirmal Raj, et al.
Published: (2024)
by: Saxena, Nirmal Raj, et al.
Published: (2024)
Multi-Objective Optimization of Consumer Group Autoscaling in Message Broker Systems
by: Landau, Diogo, et al.
Published: (2024)
by: Landau, Diogo, et al.
Published: (2024)
Fog Device-as-a-Service (FDaaS): A Framework for Service Deployment in Public Fog Environments
by: Battula, Sudheer Kumar, et al.
Published: (2023)
by: Battula, Sudheer Kumar, et al.
Published: (2023)
Integrated Sensing, Communication, and Computing: An Information-oriented Resource Transaction Mechanism
by: Chen, Ning, et al.
Published: (2024)
by: Chen, Ning, et al.
Published: (2024)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Similar Items
-
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025) -
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025) -
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
by: Qiu, Haoran, et al.
Published: (2024) -
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025) -
Mutiny! How does Kubernetes fail, and what can we do about it?
by: Barletta, Marco, et al.
Published: (2024)