Eliminating Hidden Serialization in Multi-Node Megakernel Communication
Fuente:
arXiv
Guardado en:
| Autores principales: | Oh, Byungsoo, Singh, Rachee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FlashMoE: Fast Distributed MoE in a Single Kernel
por: Aimuyo, Osayamen Jonathan, et al.
Publicado: (2025)
por: Aimuyo, Osayamen Jonathan, et al.
Publicado: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2024)
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2024)
PCCL: Photonic circuit-switched collective communication for distributed ML
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2025)
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2025)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
por: Ding, Eric, et al.
Publicado: (2026)
por: Ding, Eric, et al.
Publicado: (2026)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
por: Jin, Hongyi, et al.
Publicado: (2026)
por: Jin, Hongyi, et al.
Publicado: (2026)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
por: Sharma, Vyom, et al.
Publicado: (2026)
por: Sharma, Vyom, et al.
Publicado: (2026)
Efficient AllReduce with Stragglers
por: Devraj, Arjun, et al.
Publicado: (2025)
por: Devraj, Arjun, et al.
Publicado: (2025)
Automating Multi-Tenancy Performance Evaluation on Edge Compute Nodes
por: Georgiou, Joanna, et al.
Publicado: (2025)
por: Georgiou, Joanna, et al.
Publicado: (2025)
Short-circuiting Rings for Low-Latency AllReduce
por: Hammer, Sarah-Michelle, et al.
Publicado: (2025)
por: Hammer, Sarah-Michelle, et al.
Publicado: (2025)
Multi-Factor Trust-Driven Secure Communication Model for Cloud-Based Digital Twins
por: Saxena, Deepika, et al.
Publicado: (2026)
por: Saxena, Deepika, et al.
Publicado: (2026)
On the Universality of Round Elimination Fixed Points
por: Balliu, Alkida, et al.
Publicado: (2025)
por: Balliu, Alkida, et al.
Publicado: (2025)
Comparing Cross-Platform Performance via Node-to-Node Scaling Studies
por: Weiss, Kenneth, et al.
Publicado: (2025)
por: Weiss, Kenneth, et al.
Publicado: (2025)
Fast Switching Serial and Parallel Paradigms of SNN Inference on Multi-core Heterogeneous Neuromorphic Platform SpiNNaker2
por: Huang, Jiaxin, et al.
Publicado: (2024)
por: Huang, Jiaxin, et al.
Publicado: (2024)
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
por: Adefemi, Temitayo
Publicado: (2025)
por: Adefemi, Temitayo
Publicado: (2025)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026)
por: Kim, Taeyoon, et al.
Publicado: (2026)
Recognizing Hereditary Properties in the Presence of Byzantine Nodes
por: Cifuentes-Núñez, David, et al.
Publicado: (2023)
por: Cifuentes-Núñez, David, et al.
Publicado: (2023)
Self-healing Nodes with Adaptive Data-Sharding
por: Thakur, Ayush, et al.
Publicado: (2024)
por: Thakur, Ayush, et al.
Publicado: (2024)
Completing the Node-Averaged Complexity Landscape of LCLs on Trees
por: Balliu, Alkida, et al.
Publicado: (2024)
por: Balliu, Alkida, et al.
Publicado: (2024)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
por: Zhao, Alan, et al.
Publicado: (2026)
por: Zhao, Alan, et al.
Publicado: (2026)
Eliminating Timing Anomalies in Scheduling Periodic Segmented Self-Suspending Tasks with Release Jitter
por: Lin, Ching-Chi, et al.
Publicado: (2024)
por: Lin, Ching-Chi, et al.
Publicado: (2024)
Load Balanced Parallel Node Generation for Meshless Numerical Methods
por: Vehovar, Jon, et al.
Publicado: (2026)
por: Vehovar, Jon, et al.
Publicado: (2026)
Balancing Fixed Number of Nodes Among Multiple Fixed Clusters
por: Ranjan, Paritosh, et al.
Publicado: (2025)
por: Ranjan, Paritosh, et al.
Publicado: (2025)
Simulations between Strongly Sublinear MPC and Node-Capacitated Clique
por: Schneider, Philipp, et al.
Publicado: (2025)
por: Schneider, Philipp, et al.
Publicado: (2025)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
por: Lee, Seonho, et al.
Publicado: (2025)
por: Lee, Seonho, et al.
Publicado: (2025)
Learning Process Energy Profiles from Node-Level Power Data
por: Bader, Jonathan, et al.
Publicado: (2025)
por: Bader, Jonathan, et al.
Publicado: (2025)
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
por: Sai, Ryuichi, et al.
Publicado: (2023)
por: Sai, Ryuichi, et al.
Publicado: (2023)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
por: Ma, Xiaolong, et al.
Publicado: (2024)
por: Ma, Xiaolong, et al.
Publicado: (2024)
RAFI -- A Ray/Work Forwarding Infrastructure for Data Parallel Multi-Node/Multi-GPU Computing
por: Wald, Ingo, et al.
Publicado: (2026)
por: Wald, Ingo, et al.
Publicado: (2026)
CHIRON: Accelerating Node Synchronization without Security Trade-offs in Distributed Ledgers
por: Neiheiser, Ray, et al.
Publicado: (2024)
por: Neiheiser, Ray, et al.
Publicado: (2024)
BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain Nodes
por: Piduguralla, Manaswini, et al.
Publicado: (2026)
por: Piduguralla, Manaswini, et al.
Publicado: (2026)
Byzantine Fault Tolerant Protocols with Near-Constant Work per Node without Signatures
por: Schneider, Philipp
Publicado: (2025)
por: Schneider, Philipp
Publicado: (2025)
A Scalable State Sharing Protocol for Low-Resource Validator Nodes in Blockchain Networks
por: Hias, Ruben, et al.
Publicado: (2024)
por: Hias, Ruben, et al.
Publicado: (2024)
Advancing Blockchain Scalability: A Linear Optimization Framework for Diversified Node Allocation in Shards
por: Assmann, Björn, et al.
Publicado: (2024)
por: Assmann, Björn, et al.
Publicado: (2024)
ucTrace: A Multi-Layer Profiling Tool for UCX-driven Communication
por: Gencer, Emir, et al.
Publicado: (2026)
por: Gencer, Emir, et al.
Publicado: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
por: Chen, Jiu, et al.
Publicado: (2026)
por: Chen, Jiu, et al.
Publicado: (2026)
WANify: Gauging and Balancing Runtime WAN Bandwidth for Geo-distributed Data Analytics
por: Mohapatra, Anshuman Das, et al.
Publicado: (2025)
por: Mohapatra, Anshuman Das, et al.
Publicado: (2025)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
por: Akbari, Saman, et al.
Publicado: (2025)
por: Akbari, Saman, et al.
Publicado: (2025)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
por: Trifan, Octavian Alexandru, et al.
Publicado: (2025)
por: Trifan, Octavian Alexandru, et al.
Publicado: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
por: Qiang, Xinwei, et al.
Publicado: (2026)
por: Qiang, Xinwei, et al.
Publicado: (2026)
Ejemplares similares
-
FlashMoE: Fast Distributed MoE in a Single Kernel
por: Aimuyo, Osayamen Jonathan, et al.
Publicado: (2025) -
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2024) -
PCCL: Photonic circuit-switched collective communication for distributed ML
por: Kumar, Abhishek Vijaya, et al.
Publicado: (2025) -
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
por: Ding, Eric, et al.
Publicado: (2026) -
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)