Splitwiser: Efficient LM inference with constrained resources

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Aali, Asad, Cardoza, Adney, Capo, Melissa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913823686918144
author Aali, Asad
Cardoza, Adney
Capo, Melissa
author_facet Aali, Asad
Cardoza, Adney
Capo, Melissa
contents Efficient inference of LLMs remains a crucial challenge, with two main phases: a compute-intensive prompt computation and a memory-intensive token generation. Despite existing batching and scheduling techniques, token generation phases fail to fully utilize compute resources, especially when compared to prompt computation phases. To address these challenges, we propose Splitwiser, a methodology that splits the two phases of an LLM inference request onto the same GPU, thereby reducing overhead and improving memory access and cache utilization. By eliminating the need to transfer data across devices, Splitwiser aims to minimize network-related overheads. In this report, we describe the basic structure of our proposed pipeline while sharing preliminary results and analysis. We implement our proposed multiprocessing design on two widely-used and independent LLM architectures: Huggingface and vLLM. We open-source our code for the respective implementations: 1) Huggingface (https://github.com/asad-aali/splitwiser), and 2) vLLM (https://github.com/adney11/vllm-sysml).
format Preprint
id arxiv_https___arxiv_org_abs_2505_03763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Splitwiser: Efficient LM inference with constrained resources
Aali, Asad
Cardoza, Adney
Capo, Melissa
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
Efficient inference of LLMs remains a crucial challenge, with two main phases: a compute-intensive prompt computation and a memory-intensive token generation. Despite existing batching and scheduling techniques, token generation phases fail to fully utilize compute resources, especially when compared to prompt computation phases. To address these challenges, we propose Splitwiser, a methodology that splits the two phases of an LLM inference request onto the same GPU, thereby reducing overhead and improving memory access and cache utilization. By eliminating the need to transfer data across devices, Splitwiser aims to minimize network-related overheads. In this report, we describe the basic structure of our proposed pipeline while sharing preliminary results and analysis. We implement our proposed multiprocessing design on two widely-used and independent LLM architectures: Huggingface and vLLM. We open-source our code for the respective implementations: 1) Huggingface (https://github.com/asad-aali/splitwiser), and 2) vLLM (https://github.com/adney11/vllm-sysml).
title Splitwiser: Efficient LM inference with constrained resources
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2505.03763