DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Xin, Chen, Yuetao, Jiang, Yimin, Chen, Xing, Yan, Kun, Duan, Nan, Zhu, Yibo, Jiang, Daxin, Xu, Hong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
Echo: Simulating Distributed Training At Scale
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
Efficient Long-context Language Model Training by Core Attention Disaggregation
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
Exploiting Multicast for Accelerating Collective Communication
von: Xu, Chao, et al.
Veröffentlicht: (2026)
von: Xu, Chao, et al.
Veröffentlicht: (2026)
Exploiting the Uncertainty of the Longest Paths: Response Time Analysis for Probabilistic DAG Tasks
von: Gao, Yiyang, et al.
Veröffentlicht: (2025)
von: Gao, Yiyang, et al.
Veröffentlicht: (2025)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
von: Tan, Xin, et al.
Veröffentlicht: (2026)
von: Tan, Xin, et al.
Veröffentlicht: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Accelerating Gaussian beam tracing method with dynamic parallelism on graphics processing units
von: Sheng, Zhang, et al.
Veröffentlicht: (2025)
von: Sheng, Zhang, et al.
Veröffentlicht: (2025)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
von: Feng, Yicheng, et al.
Veröffentlicht: (2026)
von: Feng, Yicheng, et al.
Veröffentlicht: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
PICO: Accelerating All k-Core Paradigms on GPU
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
Frontier: Simulating the Next Generation of LLM Inference Systems
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
von: Chen, Kefu, et al.
Veröffentlicht: (2026)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
Enabling Dynamic Sparsity in Quantized LLM Inference
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022) -
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026) -
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024) -
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025) -
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)