vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Peiqi, Zhu, Jiangyun, Gao, Han, Zheng, Chenguang, Huang, Yongxiang, Zhou, Taichang, Yang, Ruirui, Liu, Weizhi, Chen, Weiqing, Guo, Canlin, Deng, Didan, Mo, Zifeng, Wang, Cong, Cheng, James, Wang, Roger, Liu, Hongsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910008867815424
author Yin, Peiqi
Zhu, Jiangyun
Gao, Han
Zheng, Chenguang
Huang, Yongxiang
Zhou, Taichang
Yang, Ruirui
Liu, Weizhi
Chen, Weiqing
Guo, Canlin
Deng, Didan
Mo, Zifeng
Wang, Cong
Cheng, James
Wang, Roger
Liu, Hongsheng
author_facet Yin, Peiqi
Zhu, Jiangyun
Gao, Han
Zheng, Chenguang
Huang, Yongxiang
Zhou, Taichang
Yang, Ruirui
Liu, Weizhi
Chen, Weiqing
Guo, Canlin
Deng, Didan
Mo, Zifeng
Wang, Cong
Cheng, James
Wang, Roger
Liu, Hongsheng
contents Any-to-any multimodal models that jointly handle text, images, video, and audio represent a significant advance in multimodal AI. However, their complex architectures (typically combining multiple autoregressive LLMs, diffusion transformers, and other specialized components) pose substantial challenges for efficient model serving. Existing serving systems are mainly tailored to a single paradigm, such as autoregressive LLMs for text generation or diffusion transformers for visual generation. They lack support for any-to-any pipelines that involve multiple interconnected model components. As a result, developers must manually handle cross-stage interactions, leading to huge performance degradation. We present vLLM-Omni, a fully disaggregated serving system for any-to-any models. vLLM-Omni features a novel stage abstraction that enables users to decompose complex any-to-any architectures into interconnected stages represented as a graph, and a disaggregated stage execution backend that optimizes resource utilization and throughput across stages. Each stage is independently served by an LLM or diffusion engine with per-stage request batching, flexible GPU allocation, and unified inter-stage connectors for data routing. Experimental results demonstrate that vLLM-Omni reduces job completion time (JCT) by up to 91.4% compared to baseline methods. The code is public available at https://github.com/vllm-project/vllm-omni.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02204
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
Yin, Peiqi
Zhu, Jiangyun
Gao, Han
Zheng, Chenguang
Huang, Yongxiang
Zhou, Taichang
Yang, Ruirui
Liu, Weizhi
Chen, Weiqing
Guo, Canlin
Deng, Didan
Mo, Zifeng
Wang, Cong
Cheng, James
Wang, Roger
Liu, Hongsheng
Distributed, Parallel, and Cluster Computing
Any-to-any multimodal models that jointly handle text, images, video, and audio represent a significant advance in multimodal AI. However, their complex architectures (typically combining multiple autoregressive LLMs, diffusion transformers, and other specialized components) pose substantial challenges for efficient model serving. Existing serving systems are mainly tailored to a single paradigm, such as autoregressive LLMs for text generation or diffusion transformers for visual generation. They lack support for any-to-any pipelines that involve multiple interconnected model components. As a result, developers must manually handle cross-stage interactions, leading to huge performance degradation. We present vLLM-Omni, a fully disaggregated serving system for any-to-any models. vLLM-Omni features a novel stage abstraction that enables users to decompose complex any-to-any architectures into interconnected stages represented as a graph, and a disaggregated stage execution backend that optimizes resource utilization and throughput across stages. Each stage is independently served by an LLM or diffusion engine with per-stage request batching, flexible GPU allocation, and unified inter-stage connectors for data routing. Experimental results demonstrate that vLLM-Omni reduces job completion time (JCT) by up to 91.4% compared to baseline methods. The code is public available at https://github.com/vllm-project/vllm-omni.
title vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.02204