PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914526604034048 |
|---|---|
| author | Fang, Jiarui Pan, Jinzhe Li, Aoyu Sun, Xibo Wang, Jiannan |
| author_facet | Fang, Jiarui Pan, Jinzhe Li, Aoyu Sun, Xibo Wang, Jiannan |
| contents | This paper presents PipeFusion, an innovative parallel methodology to tackle the high latency issues associated with generating high-resolution images using diffusion transformers (DiTs) models. PipeFusion partitions images into patches and the model layers across multiple GPUs. It employs a patch-level pipeline parallel strategy to orchestrate communication and computation efficiently. By capitalizing on the high similarity between inputs from successive diffusion steps, PipeFusion reuses one-step stale feature maps to provide context for the current pipeline step. This approach notably reduces communication costs compared to existing DiTs inference parallelism, including tensor parallel, sequence parallel and DistriFusion. PipeFusion enhances memory efficiency through parameter distribution across devices, ideal for large DiTs like Flux.1. Experimental results demonstrate that PipeFusion achieves state-of-the-art performance on 8$\times$L40 PCIe GPUs for Pixart, Stable-Diffusion 3, and Flux.1 models. Our source code is available at https://github.com/xdit-project/xDiT. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_14430 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference Fang, Jiarui Pan, Jinzhe Li, Aoyu Sun, Xibo Wang, Jiannan Computer Vision and Pattern Recognition Artificial Intelligence Performance This paper presents PipeFusion, an innovative parallel methodology to tackle the high latency issues associated with generating high-resolution images using diffusion transformers (DiTs) models. PipeFusion partitions images into patches and the model layers across multiple GPUs. It employs a patch-level pipeline parallel strategy to orchestrate communication and computation efficiently. By capitalizing on the high similarity between inputs from successive diffusion steps, PipeFusion reuses one-step stale feature maps to provide context for the current pipeline step. This approach notably reduces communication costs compared to existing DiTs inference parallelism, including tensor parallel, sequence parallel and DistriFusion. PipeFusion enhances memory efficiency through parameter distribution across devices, ideal for large DiTs like Flux.1. Experimental results demonstrate that PipeFusion achieves state-of-the-art performance on 8$\times$L40 PCIe GPUs for Pixart, Stable-Diffusion 3, and Flux.1 models. Our source code is available at https://github.com/xdit-project/xDiT. |
| title | PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Performance |
| url | https://arxiv.org/abs/2405.14430 |