VoxServe: Streaming-Centric Serving System for Speech Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kamahori, Keisuke, Lee, Wei-Tzu, Jha, Atindra, Kadekodi, Rohan, Wang, Stephanie, Krishnamurthy, Arvind, Kasikci, Baris
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918316437667840
author Kamahori, Keisuke
Lee, Wei-Tzu
Jha, Atindra
Kadekodi, Rohan
Wang, Stephanie
Krishnamurthy, Arvind
Kasikci, Baris
author_facet Kamahori, Keisuke
Lee, Wei-Tzu
Jha, Atindra
Kadekodi, Rohan
Wang, Stephanie
Krishnamurthy, Arvind
Kasikci, Baris
contents Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00269
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VoxServe: Streaming-Centric Serving System for Speech Language Models
Kamahori, Keisuke
Lee, Wei-Tzu
Jha, Atindra
Kadekodi, Rohan
Wang, Stephanie
Krishnamurthy, Arvind
Kasikci, Baris
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Sound
Audio and Speech Processing
Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve.
title VoxServe: Streaming-Centric Serving System for Speech Language Models
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2602.00269