SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Youxuan, Wu, Tong, Li, Shigang, Wang, Xueying, Wang, Jingjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
von: Chen, Jinfan, et al.
Veröffentlicht: (2023)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
von: Li, Shigang, et al.
Veröffentlicht: (2021)
von: Li, Shigang, et al.
Veröffentlicht: (2021)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022)
von: Li, Shigang, et al.
Veröffentlicht: (2022)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
von: Zhao, Mingjie, et al.
Veröffentlicht: (2026)
von: Zhao, Mingjie, et al.
Veröffentlicht: (2026)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
von: Murimi, Almond Kiruthu
Veröffentlicht: (2025)
von: Murimi, Almond Kiruthu
Veröffentlicht: (2025)
Connecting Large Language Model Agent to High Performance Computing Resource
von: Ma, Heng, et al.
Veröffentlicht: (2025)
von: Ma, Heng, et al.
Veröffentlicht: (2025)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
von: Bian, Zhuohang, et al.
Veröffentlicht: (2025)
von: Bian, Zhuohang, et al.
Veröffentlicht: (2025)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
von: Xiong, Yi, et al.
Veröffentlicht: (2024)
von: Xiong, Yi, et al.
Veröffentlicht: (2024)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
von: Pugachev, Sergey
Veröffentlicht: (2025)
von: Pugachev, Sergey
Veröffentlicht: (2025)
GRAIN: Exact Graph Reconstruction from Gradients
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
von: Cox, Bart, et al.
Veröffentlicht: (2022)
von: Cox, Bart, et al.
Veröffentlicht: (2022)
Roadmap for Edge AI: A Dagstuhl Perspective
von: Ding, Aaron Yi, et al.
Veröffentlicht: (2021)
von: Ding, Aaron Yi, et al.
Veröffentlicht: (2021)
Parameterizing Federated Continual Learning for Reproducible Research
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Asynchronous Byzantine Federated Learning
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
Training Diffusion Models with Federated Learning
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
von: Schroeder, Christian, et al.
Veröffentlicht: (2023)
von: Schroeder, Christian, et al.
Veröffentlicht: (2023)
MirLibSpark: A Scalable NGS Plant MicroRNA Prediction Pipeline for Multi-Library Functional Annotation
von: Wu, Chao-Jung, et al.
Veröffentlicht: (2025)
von: Wu, Chao-Jung, et al.
Veröffentlicht: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
von: Parmar, Abhinav Singh
Veröffentlicht: (2026)
von: Parmar, Abhinav Singh
Veröffentlicht: (2026)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
von: Parakhin, Vladyslav
Veröffentlicht: (2026)
von: Parakhin, Vladyslav
Veröffentlicht: (2026)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
von: Bhattacharya, Siddhartha, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Siddhartha, et al.
Veröffentlicht: (2024)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
von: Kang, Daemyung, et al.
Veröffentlicht: (2026)
von: Kang, Daemyung, et al.
Veröffentlicht: (2026)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
von: Dong, Fan, et al.
Veröffentlicht: (2023)
von: Dong, Fan, et al.
Veröffentlicht: (2023)
Uncertainty Estimation in Multi-Agent Distributed Learning for AI-Enabled Edge Devices
von: Radchenko, Gleb, et al.
Veröffentlicht: (2024)
von: Radchenko, Gleb, et al.
Veröffentlicht: (2024)
Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
von: Feng, Wenjiao, et al.
Veröffentlicht: (2025)
von: Feng, Wenjiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
von: Shi, Jinliang, et al.
Veröffentlicht: (2025) -
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024) -
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
von: Chen, Jinfan, et al.
Veröffentlicht: (2023) -
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
von: Li, Shigang, et al.
Veröffentlicht: (2021) -
Near-Optimal Sparse Allreduce for Distributed Deep Learning
von: Li, Shigang, et al.
Veröffentlicht: (2022)