VoxServe: Streaming-Centric Serving System for Speech Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kamahori, Keisuke, Lee, Wei-Tzu, Jha, Atindra, Kadekodi, Rohan, Wang, Stephanie, Krishnamurthy, Arvind, Kasikci, Baris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
FedMLAC: Mutual Learning Driven Heterogeneous Federated Audio Classification
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S
by: S., Ranjith M., et al.
Published: (2026)
by: S., Ranjith M., et al.
Published: (2026)
HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy Preservation in Multimodal Sentiment Analysis
by: Wu, Zhuojia, et al.
Published: (2024)
by: Wu, Zhuojia, et al.
Published: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
CoopASD: Cooperative Machine Anomalous Sound Detection with Privacy Concerns
by: Jiang, Anbai, et al.
Published: (2024)
by: Jiang, Anbai, et al.
Published: (2024)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
by: Zhu, Kan, et al.
Published: (2024)
by: Zhu, Kan, et al.
Published: (2024)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
by: Kamahori, Keisuke, et al.
Published: (2025)
by: Kamahori, Keisuke, et al.
Published: (2025)
Producer vs. Rapper: Who Dominates the Hip Hop Sound? A Case Study
by: Ziemer, Tim, et al.
Published: (2024)
by: Ziemer, Tim, et al.
Published: (2024)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
Offload Rethinking by Cloud Assistance for Efficient Environmental Sound Recognition on LPWANs
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)
by: Pagonas, Nikos, et al.
Published: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
by: Ye, Fanjiang, et al.
Published: (2026)
by: Ye, Fanjiang, et al.
Published: (2026)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
by: Cao, Jiahe, et al.
Published: (2026)
by: Cao, Jiahe, et al.
Published: (2026)
EdgeServe: A Streaming System for Decentralized Model Serving
by: Shaowang, Ted, et al.
Published: (2023)
by: Shaowang, Ted, et al.
Published: (2023)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
Enabling Elastic Model Serving with MultiWorld
by: Lee, Myungjin, et al.
Published: (2024)
by: Lee, Myungjin, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
by: Xiang, Yuxing, et al.
Published: (2025)
by: Xiang, Yuxing, et al.
Published: (2025)
DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
by: Strati, Foteini, et al.
Published: (2024)
by: Strati, Foteini, et al.
Published: (2024)
Version Control of Speaker Recognition Systems
by: Wang, Quan, et al.
Published: (2020)
by: Wang, Quan, et al.
Published: (2020)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
by: Du, Boxiao, et al.
Published: (2026)
by: Du, Boxiao, et al.
Published: (2026)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Similar Items
-
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026) -
FedMLAC: Mutual Learning Driven Heterogeneous Federated Audio Classification
by: Bai, Jun, et al.
Published: (2025) -
Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S
by: S., Ranjith M., et al.
Published: (2026) -
HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy Preservation in Multimodal Sentiment Analysis
by: Wu, Zhuojia, et al.
Published: (2024) -
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)