Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple
Fuente:
arXiv
Salvato in:
| Autori principali: | Georganas, Evangelos, Heinecke, Alexander, Dubey, Pradeep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
di: Georganas, Evangelos, et al.
Pubblicazione: (2023)
di: Georganas, Evangelos, et al.
Pubblicazione: (2023)
Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication
di: Brock, Benjamin, et al.
Pubblicazione: (2025)
di: Brock, Benjamin, et al.
Pubblicazione: (2025)
Towards a high-performance AI compiler with upstream MLIR
di: Golin, Renato, et al.
Pubblicazione: (2024)
di: Golin, Renato, et al.
Pubblicazione: (2024)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
di: Zhang, Dingyan, et al.
Pubblicazione: (2026)
di: Zhang, Dingyan, et al.
Pubblicazione: (2026)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
FedFT: Improving Communication Performance for Federated Learning with Frequency Space Transformation
di: Palihawadana, Chamath, et al.
Pubblicazione: (2024)
di: Palihawadana, Chamath, et al.
Pubblicazione: (2024)
ELANA: A Simple Energy and Latency Analyzer for LLMs
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
di: Zhang, Ruisi, et al.
Pubblicazione: (2024)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
di: Zhou, Zijie
Pubblicazione: (2026)
di: Zhou, Zijie
Pubblicazione: (2026)
Enhancing Communication Efficiency in FL with Adaptive Gradient Quantization and Communication Frequency Optimization
di: Tariq, Asadullah, et al.
Pubblicazione: (2025)
di: Tariq, Asadullah, et al.
Pubblicazione: (2025)
Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
di: Huang, Yafan, et al.
Pubblicazione: (2026)
di: Huang, Yafan, et al.
Pubblicazione: (2026)
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
di: Li, Qingyuan, et al.
Pubblicazione: (2025)
di: Li, Qingyuan, et al.
Pubblicazione: (2025)
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
di: Davidson, Sam, et al.
Pubblicazione: (2025)
di: Davidson, Sam, et al.
Pubblicazione: (2025)
Demystifying the Communication Characteristics for Distributed Transformer Models
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
di: Remke, Stefan, et al.
Pubblicazione: (2024)
di: Remke, Stefan, et al.
Pubblicazione: (2024)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
di: Ma, Shuang, et al.
Pubblicazione: (2026)
di: Ma, Shuang, et al.
Pubblicazione: (2026)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
di: Hwang, Changho, et al.
Pubblicazione: (2025)
di: Hwang, Changho, et al.
Pubblicazione: (2025)
iOS as Acceleration
di: Chen, Alexander K.
Pubblicazione: (2025)
di: Chen, Alexander K.
Pubblicazione: (2025)
Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding
di: Li, Chengxi, et al.
Pubblicazione: (2026)
di: Li, Chengxi, et al.
Pubblicazione: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
di: Song, Jingwei, et al.
Pubblicazione: (2025)
di: Song, Jingwei, et al.
Pubblicazione: (2025)
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
di: Liang, Feng, et al.
Pubblicazione: (2024)
di: Liang, Feng, et al.
Pubblicazione: (2024)
Fork is All You Need in Heterogeneous Systems
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
di: Liu, Man, et al.
Pubblicazione: (2026)
di: Liu, Man, et al.
Pubblicazione: (2026)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
di: Chen, Sirui, et al.
Pubblicazione: (2025)
di: Chen, Sirui, et al.
Pubblicazione: (2025)
From Data Center IoT Telemetry to Data Analytics Chatbots -- Virtual Knowledge Graph is All You Need
di: Khan, Junaid Ahmed, et al.
Pubblicazione: (2025)
di: Khan, Junaid Ahmed, et al.
Pubblicazione: (2025)
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
di: Sung, Mingyu, et al.
Pubblicazione: (2025)
di: Sung, Mingyu, et al.
Pubblicazione: (2025)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
di: Wang, Yisu, et al.
Pubblicazione: (2025)
di: Wang, Yisu, et al.
Pubblicazione: (2025)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
di: Lin, Jieke, et al.
Pubblicazione: (2025)
di: Lin, Jieke, et al.
Pubblicazione: (2025)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
di: Jiang, Xiaoxiao, et al.
Pubblicazione: (2025)
di: Jiang, Xiaoxiao, et al.
Pubblicazione: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
A Unified Convergence Analysis for Semi-Decentralized Learning: Sampled-to-Sampled vs. Sampled-to-All Communication
di: Rodio, Angelo, et al.
Pubblicazione: (2025)
di: Rodio, Angelo, et al.
Pubblicazione: (2025)
Bike Assisted Evacuation on a Line of Robots with Communication Faults
di: Jawhar, Khaled, et al.
Pubblicazione: (2023)
di: Jawhar, Khaled, et al.
Pubblicazione: (2023)
Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study
di: Wiesner, Philipp, et al.
Pubblicazione: (2026)
di: Wiesner, Philipp, et al.
Pubblicazione: (2026)
Scaling Performance of Large Language Model Pretraining
di: Interrante-Grant, Alexander, et al.
Pubblicazione: (2025)
di: Interrante-Grant, Alexander, et al.
Pubblicazione: (2025)
Ephemeral Rollups are All you Need
di: Picco, Gabriele, et al.
Pubblicazione: (2023)
di: Picco, Gabriele, et al.
Pubblicazione: (2023)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
di: Qian, Matthew, et al.
Pubblicazione: (2026)
di: Qian, Matthew, et al.
Pubblicazione: (2026)
CurvFed: Curvature-Aligned Federated Learning for Fairness without Demographics
di: Sharma, Harshit, et al.
Pubblicazione: (2024)
di: Sharma, Harshit, et al.
Pubblicazione: (2024)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
di: Metere, Alfredo
Pubblicazione: (2025)
di: Metere, Alfredo
Pubblicazione: (2025)
DGEMM on Integer Matrix Multiplication Unit
di: Ootomo, Hiroyuki, et al.
Pubblicazione: (2023)
di: Ootomo, Hiroyuki, et al.
Pubblicazione: (2023)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
di: Li, Shiju, et al.
Pubblicazione: (2025)
di: Li, Shiju, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
di: Georganas, Evangelos, et al.
Pubblicazione: (2023) -
Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication
di: Brock, Benjamin, et al.
Pubblicazione: (2025) -
Towards a high-performance AI compiler with upstream MLIR
di: Golin, Renato, et al.
Pubblicazione: (2024) -
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
di: Zhang, Dingyan, et al.
Pubblicazione: (2026) -
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)