Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework"
Fuente:
Zenodo
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902237677092864 |
|---|---|
| author | Chukkapalli, Divya Mishra, Sanjay Naik, Ganesh R |
| author_facet | Chukkapalli, Divya Mishra, Sanjay Naik, Ganesh R |
| contents | Machine-readable companion to the survey of distributed LLM-inference serving architectures: (i) a CSV taxonomy of all 56 included serving techniques with pillar and cross-pillar assignments and citation keys; (ii) the BibTeX record of every cited work (120 entries); and (iii) the stage-by-stage PRISMA-2020-style flow counts (412 -> 287 -> 198 -> 56). Released so that readers can audit pillar assignments, reproduce the reference count, and amend a classification by editing a single CSV row. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20423049 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework" Chukkapalli, Divya Mishra, Sanjay Naik, Ganesh R distributed inference LLM serving KV cache speculative decoding prefill-decode disaggregation request scheduling PRISMA taxonomy Machine-readable companion to the survey of distributed LLM-inference serving architectures: (i) a CSV taxonomy of all 56 included serving techniques with pillar and cross-pillar assignments and citation keys; (ii) the BibTeX record of every cited work (120 entries); and (iii) the stage-by-stage PRISMA-2020-style flow counts (412 -> 287 -> 198 -> 56). Released so that readers can audit pillar assignments, reproduce the reference count, and amend a classification by editing a single CSV row. |
| title | Companion artefact for "Distributed Serving Architectures for Large Language Model Inference: A Taxonomy, Quantitative Models, and Practitioner's Decision Framework" |
| topic | distributed inference LLM serving KV cache speculative decoding prefill-decode disaggregation request scheduling PRISMA taxonomy |
| url | https://doi.org/10.5281/zenodo.20423049 |