PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Sijie, Wang, Qiang, Shi, Shaohuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025)
von: Tan, Xin, et al.
Veröffentlicht: (2025)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2026)
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2026)
Partially Conditioned Patch Parallelism for Accelerated Diffusion Model Inference
von: Zhang, XiuYu, et al.
Veröffentlicht: (2024)
von: Zhang, XiuYu, et al.
Veröffentlicht: (2024)
KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Real-Time Video Generation with Pyramid Attention Broadcast
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
SwiftFusion: Scalable Sequence Parallelism for Distributed Inference of Diffusion Transformers on GPUs
von: Yang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Yang, Jiacheng, et al.
Veröffentlicht: (2026)
Task-Agnostic Federated Learning
von: Yao, Zhengtao, et al.
Veröffentlicht: (2024)
von: Yao, Zhengtao, et al.
Veröffentlicht: (2024)
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
von: Chen, Yukang, et al.
Veröffentlicht: (2026)
LAECIPS: Large Vision Model Assisted Adaptive Edge-Cloud Collaboration for IoT-based Embodied Intelligence System
von: Hu, Shijing, et al.
Veröffentlicht: (2024)
von: Hu, Shijing, et al.
Veröffentlicht: (2024)
CloudEye: A New Paradigm of Video Analysis System for Mobile Visual Scenarios
von: Cui, Huan, et al.
Veröffentlicht: (2024)
von: Cui, Huan, et al.
Veröffentlicht: (2024)
PointSplit: Towards On-device 3D Object Detection with Heterogeneous Low-power Accelerators
von: Park, Keondo, et al.
Veröffentlicht: (2025)
von: Park, Keondo, et al.
Veröffentlicht: (2025)
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
FedVSR: Towards Model-Agnostic Federated Learning in Video Super-Resolution
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2025)
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2025)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2024)
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2024)
InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
von: Wu, Zihao
Veröffentlicht: (2025)
von: Wu, Zihao
Veröffentlicht: (2025)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
von: Sela, Gur-Eyal, et al.
Veröffentlicht: (2025)
von: Sela, Gur-Eyal, et al.
Veröffentlicht: (2025)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Enabling Cross-Camera Collaboration for Video Analytics on Distributed Smart Cameras
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
von: Min, Chulhong, et al.
Veröffentlicht: (2024)
RE-POSE: Synergizing Reinforcement Learning-Based Partitioning and Offloading for Edge Object Detection
von: Shi, Jianrui, et al.
Veröffentlicht: (2025)
von: Shi, Jianrui, et al.
Veröffentlicht: (2025)
Distributed Radiance Fields for Edge Video Compression and Metaverse Integration in Autonomous Driving
von: Šlapak, Eugen, et al.
Veröffentlicht: (2024)
von: Šlapak, Eugen, et al.
Veröffentlicht: (2024)
STADI: Fine-Grained Step-Patch Diffusion Parallelism for Heterogeneous GPUs
von: Liang, Han, et al.
Veröffentlicht: (2025)
von: Liang, Han, et al.
Veröffentlicht: (2025)
Federated Learning for Diffusion Models
von: Peng, Zihao, et al.
Veröffentlicht: (2025)
von: Peng, Zihao, et al.
Veröffentlicht: (2025)
GroupNL: Low-Resource and Robust CNN Design over Cloud and Device
von: Ding, Chuntao, et al.
Veröffentlicht: (2025)
von: Ding, Chuntao, et al.
Veröffentlicht: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
Skewness-Guided Pruning of Multimodal Swin Transformers for Federated Skin Lesion Classification on Edge Devices
von: Paxton, Kuniko, et al.
Veröffentlicht: (2025)
von: Paxton, Kuniko, et al.
Veröffentlicht: (2025)
Ask the Expert: Collaborative Inference for Vision Transformers with Near-Edge Accelerators
von: Liu, Hao, et al.
Veröffentlicht: (2026)
von: Liu, Hao, et al.
Veröffentlicht: (2026)
A Scene-aware Models Adaptation Scheme for Cross-scene Online Inference on Mobile Devices
von: Li, Yunzhe, et al.
Veröffentlicht: (2024)
von: Li, Yunzhe, et al.
Veröffentlicht: (2024)
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
von: Lv, Zheqi, et al.
Veröffentlicht: (2022)
von: Lv, Zheqi, et al.
Veröffentlicht: (2022)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2024)
von: Wu, Houming, et al.
Veröffentlicht: (2024)
Data-Free Federated Class Incremental Learning with Diffusion-Based Generative Memory
von: Wang, Naibo, et al.
Veröffentlicht: (2024)
von: Wang, Naibo, et al.
Veröffentlicht: (2024)
Research on Improved U-net Based Remote Sensing Image Segmentation Algorithm
von: Yang, Qiming, et al.
Veröffentlicht: (2024)
von: Yang, Qiming, et al.
Veröffentlicht: (2024)
Federated Unsupervised Visual Representation Learning via Exploiting General Content and Personal Style
von: Yang, Yuewei, et al.
Veröffentlicht: (2022)
von: Yang, Yuewei, et al.
Veröffentlicht: (2022)
Decentralized Diffusion Models
von: McAllister, David, et al.
Veröffentlicht: (2025)
von: McAllister, David, et al.
Veröffentlicht: (2025)
Personalized Federated Fine-Tuning of Vision Foundation Models for Healthcare
von: Tupper, Adam, et al.
Veröffentlicht: (2025)
von: Tupper, Adam, et al.
Veröffentlicht: (2025)
ECORE: Energy-Conscious Optimized Routing for Deep Learning Models at the Edge
von: Alqahtani, Daghash K., et al.
Veröffentlicht: (2025)
von: Alqahtani, Daghash K., et al.
Veröffentlicht: (2025)
SemanticNN: Compressive and Error-Resilient Semantic Offloading for Extremely Weak Devices
von: Huang, Jiaming, et al.
Veröffentlicht: (2025)
von: Huang, Jiaming, et al.
Veröffentlicht: (2025)
ARIA: On the Interaction Between Architectures, Initialization and Aggregation Methods for Federated Visual Classification
von: Siomos, Vasilis, et al.
Veröffentlicht: (2023)
von: Siomos, Vasilis, et al.
Veröffentlicht: (2023)
OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery
von: Dias, Philipe, et al.
Veröffentlicht: (2024)
von: Dias, Philipe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025) -
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
von: Ramachandran, Anvitha, et al.
Veröffentlicht: (2026) -
Partially Conditioned Patch Parallelism for Accelerated Diffusion Model Inference
von: Zhang, XiuYu, et al.
Veröffentlicht: (2024) -
KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
von: Zhang, Kai, et al.
Veröffentlicht: (2025) -
Real-Time Video Generation with Pyramid Attention Broadcast
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)