MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jinghan, Jacobs, Sam Adé, Krichene, Walid, Tanaka, Masahiro, Panda, Dhabaleswar K |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
Demystifying the Communication Characteristics for Distributed Transformer Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
The Case for Co-Designing Model Architectures with Hardware
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
by: Wei, Jianyu, et al.
Published: (2024)
by: Wei, Jianyu, et al.
Published: (2024)
Deep Reinforcement Learning for Fault-Adaptive Routing in Eisenstein-Jacobi Interconnection Topologies
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
by: Charrwi, Mohammad Walid, et al.
Published: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
by: Deshmukh, Aadesh, et al.
Published: (2025)
by: Deshmukh, Aadesh, et al.
Published: (2025)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
by: Liao, Mengqi, et al.
Published: (2026)
by: Liao, Mengqi, et al.
Published: (2026)
The Power of Abstract MAC Layer: A Fault-tolerance Perspective
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Accurate Performance Predictors for Edge Computing Applications
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
Optimal-Length Labeling Schemes for Fast Deterministic Communication in Radio Networks
by: Gańczorz, Adam, et al.
Published: (2024)
by: Gańczorz, Adam, et al.
Published: (2024)
Fast Iterative Graph Computing with Updated Neighbor States
by: Zhou, Yijie, et al.
Published: (2024)
by: Zhou, Yijie, et al.
Published: (2024)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
by: Plesner, Andreas, et al.
Published: (2024)
by: Plesner, Andreas, et al.
Published: (2024)
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
by: Jacobson, John, et al.
Published: (2024)
by: Jacobson, John, et al.
Published: (2024)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
by: Pande, Sohan Kumar, et al.
Published: (2025)
by: Pande, Sohan Kumar, et al.
Published: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
by: Kumar, Madabattula Rajesh, et al.
Published: (2025)
by: Kumar, Madabattula Rajesh, et al.
Published: (2025)
Privacy-Preserving Coding Schemes for Multi-Access Distributed Computing Models
by: Sasi, Shanuja
Published: (2026)
by: Sasi, Shanuja
Published: (2026)
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
by: Xu, Zhifei, et al.
Published: (2025)
by: Xu, Zhifei, et al.
Published: (2025)
The Task Completion Problem and its Application to Crash-Resilient Computation
by: Fischer, Orr, et al.
Published: (2026)
by: Fischer, Orr, et al.
Published: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
by: Wang, Zhanghan, et al.
Published: (2025)
by: Wang, Zhanghan, et al.
Published: (2025)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
by: Kong, Jie, et al.
Published: (2025)
by: Kong, Jie, et al.
Published: (2025)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
by: Yang, Dingyu, et al.
Published: (2026)
by: Yang, Dingyu, et al.
Published: (2026)
Accurate GPU Memory Prediction for Deep Learning Jobs through Dynamic Analysis
by: Shi, Jiabo, et al.
Published: (2025)
by: Shi, Jiabo, et al.
Published: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
by: Wang, Tuowei, et al.
Published: (2025)
by: Wang, Tuowei, et al.
Published: (2025)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
CoRaiS: Lightweight Real-Time Scheduler for Multi-Edge Cooperative Computing
by: Hu, Yujiao, et al.
Published: (2024)
by: Hu, Yujiao, et al.
Published: (2024)
Similar Items
-
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024) -
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026) -
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025) -
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)