PAT: a new algorithm for all-gather and reduce-scatter operations at scale
Fuente:
arXiv
Saved in:
| Main Author: | Jeaugey, Sylvain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
by: Hu, Zhiyi, et al.
Published: (2025)
by: Hu, Zhiyi, et al.
Published: (2025)
Configurable Non-uniform All-to-all Algorithms
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Optimal, Non-pipelined Reduce-scatter and Allreduce Algorithms
by: Träff, Jesper Larsson
Published: (2024)
by: Träff, Jesper Larsson
Published: (2024)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025)
by: Yi, Jinjun, et al.
Published: (2025)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
by: Kinkead, Shannon, et al.
Published: (2026)
by: Kinkead, Shannon, et al.
Published: (2026)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
by: Tang, Xinru, et al.
Published: (2025)
by: Tang, Xinru, et al.
Published: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
Workflow decomposition algorithm for scheduling with quantum annealer-based hybrid solver
by: Kroczek, Marcin, et al.
Published: (2025)
by: Kroczek, Marcin, et al.
Published: (2025)
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
by: Frey, Matthias, et al.
Published: (2024)
by: Frey, Matthias, et al.
Published: (2024)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
by: Yang, Dingyu, et al.
Published: (2024)
by: Yang, Dingyu, et al.
Published: (2024)
Open Science Data Federation -- operation and monitoring
by: Andrijauskas, Fabio, et al.
Published: (2026)
by: Andrijauskas, Fabio, et al.
Published: (2026)
Energy-aware operation of HPC systems in Germany
by: Suarez, Estela, et al.
Published: (2024)
by: Suarez, Estela, et al.
Published: (2024)
Design of quasi phase matching crystal based on differential gray wolf algorithm
by: Chen, He, et al.
Published: (2025)
by: Chen, He, et al.
Published: (2025)
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
by: Hong, Yuxi, et al.
Published: (2024)
by: Hong, Yuxi, et al.
Published: (2024)
Fault-tolerant Reduce and Allreduce operations based on correction
by: Kuettler, Martin, et al.
Published: (2026)
by: Kuettler, Martin, et al.
Published: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Read-Modify-Writable Snapshots from Read/Write operations
by: Castañeda, Armando, et al.
Published: (2026)
by: Castañeda, Armando, et al.
Published: (2026)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
SAF: Scalable Acceleration Framework for dynamic and flexible scaling of FPGAs
by: Quraishi, Masudul Hassan, et al.
Published: (2025)
by: Quraishi, Masudul Hassan, et al.
Published: (2025)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
by: Hudson, Stephen, et al.
Published: (2024)
by: Hudson, Stephen, et al.
Published: (2024)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
by: Zhong, Yuchen, et al.
Published: (2023)
by: Zhong, Yuchen, et al.
Published: (2023)
Designing Co-operation in Systems of Hierarchical, Multi-objective Schedulers for Stream Processing
by: Dangwal, Animesh, et al.
Published: (2025)
by: Dangwal, Animesh, et al.
Published: (2025)
MOSS: A Large-scale Open Microscopic Traffic Simulation System
by: Zhang, Jun, et al.
Published: (2024)
by: Zhang, Jun, et al.
Published: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
Blockchain and Edge Computing Nexus: A Large-scale Systematic Literature Review
by: Nezami, Zeinab, et al.
Published: (2025)
by: Nezami, Zeinab, et al.
Published: (2025)
scaleTRIM: Scalable TRuncation-Based Integer Approximate Multiplier with Linearization and Compensation
by: Farahmand, Ebrahim, et al.
Published: (2023)
by: Farahmand, Ebrahim, et al.
Published: (2023)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
by: Suman, Shekhar, et al.
Published: (2026)
by: Suman, Shekhar, et al.
Published: (2026)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
by: Aguilera, Marcos K., et al.
Published: (2022)
by: Aguilera, Marcos K., et al.
Published: (2022)
Is a LOCAL algorithm computable?
by: Cruciani, Antonio, et al.
Published: (2026)
by: Cruciani, Antonio, et al.
Published: (2026)
A Hybrid Reactive-Proactive Auto-scaling Algorithm for SLA-Constrained Edge Computing
by: Gupta, Suhrid, et al.
Published: (2025)
by: Gupta, Suhrid, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
PS-WL: A Probability-Sensitive Wear Leveling scheme for SSD array scaling
by: Xu, Shuhang, et al.
Published: (2025)
by: Xu, Shuhang, et al.
Published: (2025)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
by: Song, Shengye, et al.
Published: (2025)
by: Song, Shengye, et al.
Published: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
by: Fang, Jiahao, et al.
Published: (2024)
by: Fang, Jiahao, et al.
Published: (2024)
Scrutiny new framework in integrated distributed reliable systems
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
Similar Items
-
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
by: Hu, Zhiyi, et al.
Published: (2025) -
Configurable Non-uniform All-to-all Algorithms
by: Fan, Ke, et al.
Published: (2024) -
Optimal, Non-pipelined Reduce-scatter and Allreduce Algorithms
by: Träff, Jesper Larsson
Published: (2024) -
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025) -
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
by: Kinkead, Shannon, et al.
Published: (2026)