Demystifying the Communication Characteristics for Distributed Transformer Models
Fuente:
arXiv
Saved in:
| Main Authors: | Anthony, Quentin, Michalowicz, Benjamin, Hatef, Jacob, Xu, Lang, Abduljabbar, Mustafa, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
The Case for Co-Designing Model Architectures with Hardware
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
by: Zhu, Zhuoran, et al.
Published: (2025)
by: Zhu, Zhuoran, et al.
Published: (2025)
Comparative Study of Large Language Model Architectures on Frontier
by: Yin, Junqi, et al.
Published: (2024)
by: Yin, Junqi, et al.
Published: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)
by: Shyam, Vasu, et al.
Published: (2026)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
Distributed Resource Selection for Self-Organising Cloud-Edge Systems
by: Renau, Quentin, et al.
Published: (2025)
by: Renau, Quentin, et al.
Published: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
by: Pan, Qilong, et al.
Published: (2025)
by: Pan, Qilong, et al.
Published: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Galvatron: Automatic Distributed Training for Large Transformer Models
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
by: Hu, Zhiyi, et al.
Published: (2025)
by: Hu, Zhiyi, et al.
Published: (2025)
Pilotfish: Distributed Execution for Scalable Blockchains
by: Kniep, Quentin, et al.
Published: (2024)
by: Kniep, Quentin, et al.
Published: (2024)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
by: Chen, Xu, et al.
Published: (2024)
by: Chen, Xu, et al.
Published: (2024)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
by: Xu, Guanbin, et al.
Published: (2026)
by: Xu, Guanbin, et al.
Published: (2026)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
by: Ma, Bin, et al.
Published: (2026)
by: Ma, Bin, et al.
Published: (2026)
The R(1)W(1) Communication Model for Self-Stabilizing Distributed Algorithms
by: Kakugawa, Hirotsugu, et al.
Published: (2025)
by: Kakugawa, Hirotsugu, et al.
Published: (2025)
Distributed Ranges: A Model for Distributed Data Structures, Algorithms, and Views
by: Brock, Benjamin, et al.
Published: (2024)
by: Brock, Benjamin, et al.
Published: (2024)
Demystifying Object-based Big Data Storage Systems
by: Mondal, Anindita Sarkar, et al.
Published: (2024)
by: Mondal, Anindita Sarkar, et al.
Published: (2024)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
by: Akbari, Negin, et al.
Published: (2025)
by: Akbari, Negin, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
Distributed Statistical Zero-Knowledge Proofs via Sumcheck
by: Jauregui, Benjamin, et al.
Published: (2026)
by: Jauregui, Benjamin, et al.
Published: (2026)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026)
by: Maurya, Avinash, et al.
Published: (2026)
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
by: Zhu, Ying, et al.
Published: (2025)
by: Zhu, Ying, et al.
Published: (2025)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
by: Wahlgren, Jacob, et al.
Published: (2026)
by: Wahlgren, Jacob, et al.
Published: (2026)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
by: Sun, Ao, et al.
Published: (2025)
by: Sun, Ao, et al.
Published: (2025)
Dynamic Contract Analysis for Parallel Programming Models
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
by: Oraji, Yussur Mustafa, et al.
Published: (2026)
A HPX Communication Benchmark: Distributed FFT using Collectives
by: Strack, Alexander, et al.
Published: (2025)
by: Strack, Alexander, et al.
Published: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
BCL: A Cross-Platform Distributed Container Library
by: Brock, Benjamin, et al.
Published: (2018)
by: Brock, Benjamin, et al.
Published: (2018)
COoL-TEE: Client-TEE Collaboration for Resilient Distributed Search
by: Bettinger, Matthieu, et al.
Published: (2025)
by: Bettinger, Matthieu, et al.
Published: (2025)
Strong and Hiding Distributed Certification of Bipartiteness
by: Jauregui, Benjamin, et al.
Published: (2025)
by: Jauregui, Benjamin, et al.
Published: (2025)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
Similar Items
-
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024) -
The Case for Co-Designing Model Architectures with Hardware
by: Anthony, Quentin, et al.
Published: (2024) -
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024) -
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)