Beyond the Buzz: A Pragmatic Take on Inference Disaggregation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mitra, Tiyasa, Borkar, Ritika, Bhatia, Nidhi, Matas, Ramon, Raj, Shivam, Mudigere, Dheevatsa, Zhao, Ritchie, Golub, Maximilian, Dutta, Arpan, Madduri, Sailaja, Jani, Dharmesh, Pharris, Brian, Rouhani, Bita Darvish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912416073252864
author Mitra, Tiyasa
Borkar, Ritika
Bhatia, Nidhi
Matas, Ramon
Raj, Shivam
Mudigere, Dheevatsa
Zhao, Ritchie
Golub, Maximilian
Dutta, Arpan
Madduri, Sailaja
Jani, Dharmesh
Pharris, Brian
Rouhani, Bita Darvish
author_facet Mitra, Tiyasa
Borkar, Ritika
Bhatia, Nidhi
Matas, Ramon
Raj, Shivam
Mudigere, Dheevatsa
Zhao, Ritchie
Golub, Maximilian
Dutta, Arpan
Madduri, Sailaja
Jani, Dharmesh
Pharris, Brian
Rouhani, Bita Darvish
contents As inference scales to multi-node deployments, disaggregation - splitting inference into distinct phases - offers a promising path to improving the throughput-interactivity Pareto frontier. Despite growing enthusiasm and a surge of open-source efforts, practical deployment of disaggregated serving remains limited due to the complexity of the optimization search space and system-level coordination. In this paper, we present the first systematic study of disaggregated inference at scale, evaluating hundreds of thousands of design points across diverse workloads and hardware configurations. We find that disaggregation is most effective for prefill-heavy traffic patterns and larger models. Our results highlight the critical role of dynamic rate matching and elastic scaling in achieving Pareto-optimal performance. Our findings offer actionable insights for efficient disaggregated deployments to navigate the trade-off between system throughput and interactivity.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
Mitra, Tiyasa
Borkar, Ritika
Bhatia, Nidhi
Matas, Ramon
Raj, Shivam
Mudigere, Dheevatsa
Zhao, Ritchie
Golub, Maximilian
Dutta, Arpan
Madduri, Sailaja
Jani, Dharmesh
Pharris, Brian
Rouhani, Bita Darvish
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
As inference scales to multi-node deployments, disaggregation - splitting inference into distinct phases - offers a promising path to improving the throughput-interactivity Pareto frontier. Despite growing enthusiasm and a surge of open-source efforts, practical deployment of disaggregated serving remains limited due to the complexity of the optimization search space and system-level coordination. In this paper, we present the first systematic study of disaggregated inference at scale, evaluating hundreds of thousands of design points across diverse workloads and hardware configurations. We find that disaggregation is most effective for prefill-heavy traffic patterns and larger models. Our results highlight the critical role of dynamic rate matching and elastic scaling in achieving Pareto-optimal performance. Our findings offer actionable insights for efficient disaggregated deployments to navigate the trade-off between system throughput and interactivity.
title Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2506.05508