LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Lingyun, Li, Suyi, Feng, Tianyu, Jiang, Xiaoxiao, Di, Zhipeng, Lu, Weiyi, Liu, Kan, Yu, Yinghao, Lan, Tao, Yang, Guodong, Qu, Lin, Zhang, Liping, Wang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SwiftDiffusion: Efficient Diffusion Model Serving with Add-on Modules
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
von: Yang, Qizheng, et al.
Veröffentlicht: (2025)
von: Yang, Qizheng, et al.
Veröffentlicht: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
Taming the Memory Footprint Crisis: System Design for Production Diffusion LLM Serving
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
von: Wu, Jing, et al.
Veröffentlicht: (2025)
von: Wu, Jing, et al.
Veröffentlicht: (2025)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
von: Moraru, Maxim, et al.
Veröffentlicht: (2025)
von: Moraru, Maxim, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
von: Li, Cong, et al.
Veröffentlicht: (2025)
von: Li, Cong, et al.
Veröffentlicht: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
MoLink: Distributed and Efficient Serving Framework for Large Models
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SwiftDiffusion: Efficient Diffusion Model Serving with Add-on Modules
von: Li, Suyi, et al.
Veröffentlicht: (2024) -
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025) -
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024) -
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025) -
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
von: Yang, Qizheng, et al.
Veröffentlicht: (2025)