DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jinxiao, Xu, Yunpu, Wu, Xiyong, Dong, Runmin, Cheng, Shenggan, Zhao, Yi, Chen, Mengxuan, Zheng, Qinrui, Liu, Jianting, Fu, Haohuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025)
von: Tan, Xin, et al.
Veröffentlicht: (2025)
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)
von: Brown, Nick, et al.
Veröffentlicht: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
von: Tharwani, Jay, et al.
Veröffentlicht: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics
von: Austin, Allison, et al.
Veröffentlicht: (2026)
von: Austin, Allison, et al.
Veröffentlicht: (2026)
FedHC: A Hierarchical Clustered Federated Learning Framework for Satellite Networks
von: Liu, Zhuocheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhuocheng, et al.
Veröffentlicht: (2025)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
von: Shan, Baodi, et al.
Veröffentlicht: (2025)
Towards Energy Efficient Co-Scheduling in HPC
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
von: Iserte, Sergio, et al.
Veröffentlicht: (2025)
von: Iserte, Sergio, et al.
Veröffentlicht: (2025)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows
von: Etz, Brian D., et al.
Veröffentlicht: (2025)
von: Etz, Brian D., et al.
Veröffentlicht: (2025)
Efficient Column-Wise N:M Pruning on RISC-V CPU
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
UNR: Unified Notifiable RMA Library for HPC
von: Feng, Guangnan, et al.
Veröffentlicht: (2024)
von: Feng, Guangnan, et al.
Veröffentlicht: (2024)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
von: Georganas, Evangelos, et al.
Veröffentlicht: (2023)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2023)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
Advances in Semantic Patching for HPC-oriented Refactorings with Coccinelle
von: Martone, Michele, et al.
Veröffentlicht: (2025)
von: Martone, Michele, et al.
Veröffentlicht: (2025)
Attack Graph Generation on HPC Clusters
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
von: Jaghouar, Sami, et al.
Veröffentlicht: (2024)
von: Jaghouar, Sami, et al.
Veröffentlicht: (2024)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
DiReDi: Distillation and Reverse Distillation for AIoT Applications
von: Sun, Chen, et al.
Veröffentlicht: (2024)
von: Sun, Chen, et al.
Veröffentlicht: (2024)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025) -
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024) -
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025) -
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026) -
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
von: Brown, Nick, et al.
Veröffentlicht: (2024)