VoxServe: Streaming-Centric Serving System for Speech Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918316437667840 |
|---|---|
| author | Kamahori, Keisuke Lee, Wei-Tzu Jha, Atindra Kadekodi, Rohan Wang, Stephanie Krishnamurthy, Arvind Kasikci, Baris |
| author_facet | Kamahori, Keisuke Lee, Wei-Tzu Jha, Atindra Kadekodi, Rohan Wang, Stephanie Krishnamurthy, Arvind Kasikci, Baris |
| contents | Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_00269 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VoxServe: Streaming-Centric Serving System for Speech Language Models Kamahori, Keisuke Lee, Wei-Tzu Jha, Atindra Kadekodi, Rohan Wang, Stephanie Krishnamurthy, Arvind Kasikci, Baris Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing Sound Audio and Speech Processing Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existing systems fall short of supporting diverse models flexibly and efficiently. We present VoxServe, a unified serving system for SpeechLMs that optimizes streaming performance. VoxServe introduces a model-execution abstraction that decouples model architecture from system-level optimizations, thereby enabling support for diverse SpeechLM architectures within a single framework. Building on this abstraction, VoxServe implements streaming-aware scheduling and an asynchronous inference pipeline to improve end-to-end efficiency. Evaluations across multiple modern SpeechLMs show that VoxServe achieves 10-20x higher throughput than existing implementations at comparable latency while maintaining high streaming viability. The code of VoxServe is available at https://github.com/vox-serve/vox-serve. |
| title | VoxServe: Streaming-Centric Serving System for Speech Language Models |
| topic | Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2602.00269 |