Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework"

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chukkapalli, Divya, Mishra, Sanjay, Naik, Ganesh R
Format: Recurso digital
Publié: Zenodo 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866902237677092864
author Chukkapalli, Divya
Mishra, Sanjay
Naik, Ganesh R
author_facet Chukkapalli, Divya
Mishra, Sanjay
Naik, Ganesh R
contents Machine-readable companion to the survey of distributed LLM-inference serving architectures: (i) a CSV taxonomy of all 56 included serving techniques with pillar and cross-pillar assignments and citation keys; (ii) the BibTeX record of every cited work (120 entries); and (iii) the stage-by-stage PRISMA-2020-style flow counts (412 -> 287 -> 198 -> 56). Released so that readers can audit pillar assignments, reproduce the reference count, and amend a classification by editing a single CSV row.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20423049
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework"
Chukkapalli, Divya
Mishra, Sanjay
Naik, Ganesh R
distributed inference
LLM serving
KV cache
speculative decoding
prefill-decode disaggregation
request scheduling
PRISMA
taxonomy
Machine-readable companion to the survey of distributed LLM-inference serving architectures: (i) a CSV taxonomy of all 56 included serving techniques with pillar and cross-pillar assignments and citation keys; (ii) the BibTeX record of every cited work (120 entries); and (iii) the stage-by-stage PRISMA-2020-style flow counts (412 -> 287 -> 198 -> 56). Released so that readers can audit pillar assignments, reproduce the reference count, and amend a classification by editing a single CSV row.
title Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework"
topic distributed inference
LLM serving
KV cache
speculative decoding
prefill-decode disaggregation
request scheduling
PRISMA
taxonomy
url https://doi.org/10.5281/zenodo.20423049