Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Ruilong, Li, Xinjiao, Wang, Yisu, Chen, Xinyu, Kutscher, Dirk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025)
by: Guo, Jihu, et al.
Published: (2025)
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
by: Shi, Ziji, et al.
Published: (2023)
by: Shi, Ziji, et al.
Published: (2023)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
by: Los, Denis, et al.
Published: (2025)
by: Los, Denis, et al.
Published: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
Research on Model Parallelism and Data Parallelism Optimization Methods in Large Language Model-Based Recommendation Systems
by: Yang, Haowei, et al.
Published: (2025)
by: Yang, Haowei, et al.
Published: (2025)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
by: Barron, Ryan, et al.
Published: (2024)
by: Barron, Ryan, et al.
Published: (2024)
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
by: Wang, Xiaoye
Published: (2024)
by: Wang, Xiaoye
Published: (2024)
High-Throughput LLM inference on Heterogeneous Clusters
by: Xiong, Yi, et al.
Published: (2025)
by: Xiong, Yi, et al.
Published: (2025)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
by: Fang, Jiarui, et al.
Published: (2024)
by: Fang, Jiarui, et al.
Published: (2024)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
by: Wang, Li, et al.
Published: (2024)
by: Wang, Li, et al.
Published: (2024)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
by: Hwang, Changho, et al.
Published: (2025)
by: Hwang, Changho, et al.
Published: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
by: Li, Wanqian, et al.
Published: (2026)
by: Li, Wanqian, et al.
Published: (2026)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Joint Resource Optimization, Computation Offloading and Resource Slicing for Multi-Edge Traffic-Cognitive Networks
by: Xiaoyang, Ting, et al.
Published: (2024)
by: Xiaoyang, Ting, et al.
Published: (2024)
FedDCT: A Dynamic Cross-Tier Federated Learning Framework in Wireless Networks
by: Xian, Youquan, et al.
Published: (2023)
by: Xian, Youquan, et al.
Published: (2023)
Janus: Collaborative Vision Transformer Under Dynamic Network Environment
by: Jiang, Linyi, et al.
Published: (2025)
by: Jiang, Linyi, et al.
Published: (2025)
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
by: Jin, Jianli, et al.
Published: (2025)
by: Jin, Jianli, et al.
Published: (2025)
From Coordinate Matching to Structural Alignment: Rethinking Prototype Alignment in Heterogeneous Federated Learning
by: Wu, Xinghao, et al.
Published: (2026)
by: Wu, Xinghao, et al.
Published: (2026)
Para-B&B: Load-Balanced Deterministic Parallelization of Solving MIP
by: Zhang, Jinyu, et al.
Published: (2026)
by: Zhang, Jinyu, et al.
Published: (2026)
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
by: Kilic, Ozgur O., et al.
Published: (2025)
by: Kilic, Ozgur O., et al.
Published: (2025)
High-Performance Parallel Optimization of the Fish School Behaviour on the Setonix Platform Using OpenMP
by: Wang, Haitian, et al.
Published: (2025)
by: Wang, Haitian, et al.
Published: (2025)
FreeRide: Harvesting Bubbles in Pipeline Parallelism
by: Zhang, Jiashu, et al.
Published: (2024)
by: Zhang, Jiashu, et al.
Published: (2024)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
by: Wang, Shiju, et al.
Published: (2025)
by: Wang, Shiju, et al.
Published: (2025)
A Model Aware AIGC Task Offloading Algorithm in IIoT Edge Computing
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Collaborative Split Federated Learning with Parallel Training and Aggregation
by: Papageorgiou, Yiannis, et al.
Published: (2025)
by: Papageorgiou, Yiannis, et al.
Published: (2025)
Can Large Language Models Write Parallel Code?
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
Deploying Graph Neural Networks in Wireless Networks: A Link Stability Viewpoint
by: Li, Jun, et al.
Published: (2024)
by: Li, Jun, et al.
Published: (2024)
ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and Translation
by: Kaplan, Erel, et al.
Published: (2026)
by: Kaplan, Erel, et al.
Published: (2026)
TimelyFreeze: Adaptive Parameter Freezing Mechanism for Pipeline Parallelism
by: Cho, Seonghye, et al.
Published: (2026)
by: Cho, Seonghye, et al.
Published: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
by: Li, Rongzhi, et al.
Published: (2025)
by: Li, Rongzhi, et al.
Published: (2025)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
RLHFless: Serverless Computing for Efficient RLHF
by: Wei, Rui, et al.
Published: (2026)
by: Wei, Rui, et al.
Published: (2026)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
by: Zhang, Ruisi, et al.
Published: (2024)
by: Zhang, Ruisi, et al.
Published: (2024)
Similar Items
-
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025) -
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025) -
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
by: Wang, Peiran, et al.
Published: (2025) -
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025) -
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
by: Shi, Ziji, et al.
Published: (2023)