Fused3S: Fast Sparse Attention on Tensor Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zitong, Chandramowlishwaran, Aparna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
von: Cai, Weilin, et al.
Veröffentlicht: (2025)
von: Cai, Weilin, et al.
Veröffentlicht: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Efficient Long-context Language Model Training by Core Attention Disaggregation
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
Morphling: Fast, Fused, and Flexible GNN Training at Scale
von: Anubhab, et al.
Veröffentlicht: (2025)
von: Anubhab, et al.
Veröffentlicht: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
Averaging Rate Scheduler for Decentralized Learning on Heterogeneous Data
von: Aketi, Sai Aparna, et al.
Veröffentlicht: (2024)
von: Aketi, Sai Aparna, et al.
Veröffentlicht: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
Overcoming Challenges of Partial Client Participation in Federated Learning : A Comprehensive Review
von: Sen, Mrinmay, et al.
Veröffentlicht: (2025)
von: Sen, Mrinmay, et al.
Veröffentlicht: (2025)
Homogenizing Non-IID datasets via In-Distribution Knowledge Distillation for Decentralized Learning
von: Ravikumar, Deepak, et al.
Veröffentlicht: (2023)
von: Ravikumar, Deepak, et al.
Veröffentlicht: (2023)
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
von: Shi, Ziji, et al.
Veröffentlicht: (2023)
von: Shi, Ziji, et al.
Veröffentlicht: (2023)
ShardTensor: Domain Parallelism for Scientific Machine Learning
von: Adams, Corey, et al.
Veröffentlicht: (2026)
von: Adams, Corey, et al.
Veröffentlicht: (2026)
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
von: Zhao, Pinxue, et al.
Veröffentlicht: (2024)
von: Zhao, Pinxue, et al.
Veröffentlicht: (2024)
TensorSocket: Shared Data Loading for Deep Learning Training
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
von: Robroek, Ties, et al.
Veröffentlicht: (2024)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
Fast Decentralized Gradient Tracking for Federated Minimax Optimization with Local Updates
von: Li, Chris Junchi
Veröffentlicht: (2024)
von: Li, Chris Junchi
Veröffentlicht: (2024)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
Federated LoRA with Sparse Communication
von: Kuo, Kevin, et al.
Veröffentlicht: (2024)
von: Kuo, Kevin, et al.
Veröffentlicht: (2024)
Stochastic Sparse Attention for Memory-Bound Inference
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
Tackling Resource-Constrained and Data-Heterogeneity in Federated Learning with Double-Weight Sparse Pack
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
Is Flash Attention Stable?
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
Learnable Sparse Customization in Heterogeneous Edge Computing
von: Xue, Jingjing, et al.
Veröffentlicht: (2024)
von: Xue, Jingjing, et al.
Veröffentlicht: (2024)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
von: GU, Qiqi, et al.
Veröffentlicht: (2025)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
von: Afroz, Sabiha, et al.
Veröffentlicht: (2025)
von: Afroz, Sabiha, et al.
Veröffentlicht: (2025)
Delta Tensor: Efficient Vector and Tensor Storage in Delta Lake
von: Bao, Zhiwei, et al.
Veröffentlicht: (2024)
von: Bao, Zhiwei, et al.
Veröffentlicht: (2024)
Towards Communication-efficient Federated Learning via Sparse and Aligned Adaptive Optimization
von: Deng, Xiumei, et al.
Veröffentlicht: (2024)
von: Deng, Xiumei, et al.
Veröffentlicht: (2024)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
von: Xia, Yufan, et al.
Veröffentlicht: (2024)
von: Xia, Yufan, et al.
Veröffentlicht: (2024)
FedNS: A Fast Sketching Newton-Type Algorithm for Federated Learning
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Fed-GAME: Personalized Federated Learning with Graph Attention Mixture-of-Experts For Time-Series Forecasting
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
von: Shen, Ao, et al.
Veröffentlicht: (2024)
von: Shen, Ao, et al.
Veröffentlicht: (2024)
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
von: Wu, Chenpeng, et al.
Veröffentlicht: (2025)
von: Wu, Chenpeng, et al.
Veröffentlicht: (2025)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
von: Yu, Zhongming, et al.
Veröffentlicht: (2024)
von: Yu, Zhongming, et al.
Veröffentlicht: (2024)
Optimizing the Optimal Weighted Average: Efficient Distributed Sparse Classification
von: Lu, Fred, et al.
Veröffentlicht: (2024)
von: Lu, Fred, et al.
Veröffentlicht: (2024)
FedGAT: A Privacy-Preserving Federated Approximation Algorithm for Graph Attention Networks
von: Ambekar, Siddharth, et al.
Veröffentlicht: (2024)
von: Ambekar, Siddharth, et al.
Veröffentlicht: (2024)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
von: Jiang, Yinsicheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
von: Cai, Weilin, et al.
Veröffentlicht: (2025) -
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024) -
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025) -
Efficient Long-context Language Model Training by Core Attention Disaggregation
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025) -
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)