Niyama : Breaking the Silos of LLM Inference Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goel, Kanishk, Mohan, Jayashree, Kwatra, Nipun, Anupindi, Ravi Shreyas, Ramjee, Ramachandran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912299515641856
author Goel, Kanishk
Mohan, Jayashree
Kwatra, Nipun
Anupindi, Ravi Shreyas
Ramjee, Ramachandran
author_facet Goel, Kanishk
Mohan, Jayashree
Kwatra, Nipun
Anupindi, Ravi Shreyas
Ramjee, Ramachandran
contents The widespread adoption of Large Language Models (LLMs) has enabled diverse applications with very different latency requirements. Existing LLM serving frameworks rely on siloed infrastructure with coarse-grained workload segregation -- interactive and batch -- leading to inefficient resource utilization and limited support for fine-grained Quality-of-Service (QoS) differentiation. This results in operational inefficiencies, over-provisioning and poor load management during traffic surges. We present Niyama, a novel QoS-driven inference serving system that enables efficient co-scheduling of diverse workloads on shared infrastructure. Niyama introduces fine-grained QoS classification allowing applications to specify precise latency requirements, and dynamically adapts scheduling decisions based on real-time system state. Leveraging the predictable execution characteristics of LLM inference, Niyama implements a dynamic chunking mechanism to improve overall throughput while maintaining strict QoS guarantees. Additionally, Niyama employs a hybrid prioritization policy that balances fairness and efficiency, and employs selective request relegation that enables graceful service degradation during overload conditions. Our evaluation demonstrates that Niyama increases serving capacity by 32% compared to current siloed deployments, while maintaining QoS guarantees. Notably, under extreme load, our system reduces SLO violations by an order of magnitude compared to current strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Niyama : Breaking the Silos of LLM Inference Serving
Goel, Kanishk
Mohan, Jayashree
Kwatra, Nipun
Anupindi, Ravi Shreyas
Ramjee, Ramachandran
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
The widespread adoption of Large Language Models (LLMs) has enabled diverse applications with very different latency requirements. Existing LLM serving frameworks rely on siloed infrastructure with coarse-grained workload segregation -- interactive and batch -- leading to inefficient resource utilization and limited support for fine-grained Quality-of-Service (QoS) differentiation. This results in operational inefficiencies, over-provisioning and poor load management during traffic surges. We present Niyama, a novel QoS-driven inference serving system that enables efficient co-scheduling of diverse workloads on shared infrastructure. Niyama introduces fine-grained QoS classification allowing applications to specify precise latency requirements, and dynamically adapts scheduling decisions based on real-time system state. Leveraging the predictable execution characteristics of LLM inference, Niyama implements a dynamic chunking mechanism to improve overall throughput while maintaining strict QoS guarantees. Additionally, Niyama employs a hybrid prioritization policy that balances fairness and efficiency, and employs selective request relegation that enables graceful service degradation during overload conditions. Our evaluation demonstrates that Niyama increases serving capacity by 32% compared to current siloed deployments, while maintaining QoS guarantees. Notably, under extreme load, our system reduces SLO violations by an order of magnitude compared to current strategies.
title Niyama : Breaking the Silos of LLM Inference Serving
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.22562