HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Qizheng, Chen, Tung-I, Zhao, Siyu, Sitaraman, Ramesh K., Guan, Hui
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912803616456704
author Yang, Qizheng
Chen, Tung-I
Zhao, Siyu
Sitaraman, Ramesh K.
Guan, Hui
author_facet Yang, Qizheng
Chen, Tung-I
Zhao, Siyu
Sitaraman, Ramesh K.
Guan, Hui
contents Text-to-image diffusion models have achieved remarkable visual quality but incur high computational costs, making latency-aware, scalable deployment challenging. To address this, we advocate a hybrid architecture that achieves query awareness when serving diffusion models. Unlike existing query-aware serving systems that cascade lightweight and heavyweight models with a fixed configuration, our hybrid architecture first routes each query directly to a suitable model variant, then reroutes it to a cascaded heavyweight model only if necessary. We theoretically analyze conditions for the hybrid architecture to outperform non-hybrid alternatives in latency and response quality. Building on this architecture, we design HADIS, a hybrid serving system for latency-aware diffusion models that jointly optimizes cascade model selection, query routing, and resource allocation. To reduce the complexity of resource management, HADIS uses an offline profiling phase to produce a Pareto-optimal cascade configuration table. At runtime, HADIS selects the best cascade configuration and GPU allocation given latency and workload constraints. Empirical evaluations on real-world traces demonstrate that HADIS improves response quality by up to 35% while reducing latency violation rates by 2.7-45$\times$ compared to state-of-the-art model serving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
Yang, Qizheng
Chen, Tung-I
Zhao, Siyu
Sitaraman, Ramesh K.
Guan, Hui
Distributed, Parallel, and Cluster Computing
Text-to-image diffusion models have achieved remarkable visual quality but incur high computational costs, making latency-aware, scalable deployment challenging. To address this, we advocate a hybrid architecture that achieves query awareness when serving diffusion models. Unlike existing query-aware serving systems that cascade lightweight and heavyweight models with a fixed configuration, our hybrid architecture first routes each query directly to a suitable model variant, then reroutes it to a cascaded heavyweight model only if necessary. We theoretically analyze conditions for the hybrid architecture to outperform non-hybrid alternatives in latency and response quality. Building on this architecture, we design HADIS, a hybrid serving system for latency-aware diffusion models that jointly optimizes cascade model selection, query routing, and resource allocation. To reduce the complexity of resource management, HADIS uses an offline profiling phase to produce a Pareto-optimal cascade configuration table. At runtime, HADIS selects the best cascade configuration and GPU allocation given latency and workload constraints. Empirical evaluations on real-world traces demonstrate that HADIS improves response quality by up to 35% while reducing latency violation rates by 2.7-45$\times$ compared to state-of-the-art model serving systems.
title HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.00642