ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
Fuente:
arXiv
Saved in:
| Main Author: | Elango, Venmugil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024)
by: Elango, Venmugil
Published: (2024)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
by: Coquelin, Daniel, et al.
Published: (2024)
by: Coquelin, Daniel, et al.
Published: (2024)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
by: Yi, Kai, et al.
Published: (2024)
by: Yi, Kai, et al.
Published: (2024)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
by: Li, Dacheng, et al.
Published: (2023)
by: Li, Dacheng, et al.
Published: (2023)
Distributed Low-Communication Training with Decoupled Momentum Optimization
by: Nedelkoski, Sasho, et al.
Published: (2025)
by: Nedelkoski, Sasho, et al.
Published: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
by: Deng, Xiumei, et al.
Published: (2025)
by: Deng, Xiumei, et al.
Published: (2025)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
by: Heinsen, Franz A., et al.
Published: (2026)
by: Heinsen, Franz A., et al.
Published: (2026)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
by: Holen, Martin, et al.
Published: (2023)
by: Holen, Martin, et al.
Published: (2023)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
by: Oakley, Joe, et al.
Published: (2024)
by: Oakley, Joe, et al.
Published: (2024)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Efficient Resource Scheduling for Distributed Infrastructures Using Negotiation Capabilities
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Communication-Efficient Split Learning via Adaptive Feature-Wise Compression
by: Oh, Yongjeong, et al.
Published: (2023)
by: Oh, Yongjeong, et al.
Published: (2023)
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
by: Shen, Yuanzhe, et al.
Published: (2025)
by: Shen, Yuanzhe, et al.
Published: (2025)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
by: Zhang, Biyao, et al.
Published: (2025)
by: Zhang, Biyao, et al.
Published: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
Fed-Sophia: A Communication-Efficient Second-Order Federated Learning Algorithm
by: Elbakary, Ahmed, et al.
Published: (2024)
by: Elbakary, Ahmed, et al.
Published: (2024)
Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition
by: Liu, Ruyue, et al.
Published: (2024)
by: Liu, Ruyue, et al.
Published: (2024)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
by: Moore, Hayden, et al.
Published: (2025)
by: Moore, Hayden, et al.
Published: (2025)
E-3SFC: Communication-Efficient Federated Learning with Double-way Features Synthesizing
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
SpaFL: Communication-Efficient Federated Learning with Sparse Models and Low computational Overhead
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural Networks
by: Guliyev, Rustam, et al.
Published: (2024)
by: Guliyev, Rustam, et al.
Published: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026)
by: Wu, Hanjiang, et al.
Published: (2026)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Communication-Efficient Federated Learning for LEO Satellite Networks Integrated with HAPs Using Hybrid NOMA-OFDM
by: Elmahallawy, Mohamed, et al.
Published: (2024)
by: Elmahallawy, Mohamed, et al.
Published: (2024)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
by: Wei, Cunyang, et al.
Published: (2026)
by: Wei, Cunyang, et al.
Published: (2026)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
by: Cao, Linxiao, et al.
Published: (2024)
by: Cao, Linxiao, et al.
Published: (2024)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
CommunityAI: Towards Community-based Federated Learning
by: Murturi, Ilir, et al.
Published: (2023)
by: Murturi, Ilir, et al.
Published: (2023)
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
by: Dai, Huangliang, et al.
Published: (2025)
by: Dai, Huangliang, et al.
Published: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)
by: Deshmukh, Dhruv, et al.
Published: (2025)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
by: Chamroukhi, Faïcel, et al.
Published: (2023)
by: Chamroukhi, Faïcel, et al.
Published: (2023)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
On the Fragility of Data Attribution When Learning Is Distributed
by: Gao, Xian, et al.
Published: (2026)
by: Gao, Xian, et al.
Published: (2026)
Measuring Heterogeneity in Machine Learning with Distributed Energy Distance
by: Fan, Mengchen, et al.
Published: (2025)
by: Fan, Mengchen, et al.
Published: (2025)
Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
by: Choukroun, Yoni, et al.
Published: (2024)
by: Choukroun, Yoni, et al.
Published: (2024)
Similar Items
-
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024) -
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
by: Coquelin, Daniel, et al.
Published: (2024) -
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
by: Yi, Kai, et al.
Published: (2024) -
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
by: Li, Dacheng, et al.
Published: (2023) -
Distributed Low-Communication Training with Decoupled Momentum Optimization
by: Nedelkoski, Sasho, et al.
Published: (2025)