Accelerating Sparse DNNs Based on Tiled GEMM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Cong, Xue, Fengchen, Leng, Jingwen, Qiu, Yuxian, Guan, Yue, Cui, Weihao, Chen, Quan, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
von: Zhou, Yangjie, et al.
Veröffentlicht: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
von: Huang, Ziyu, et al.
Veröffentlicht: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
von: Xu, Ao, et al.
Veröffentlicht: (2025)
von: Xu, Ao, et al.
Veröffentlicht: (2025)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
von: Xue, Weicheng, et al.
Veröffentlicht: (2025)
von: Xue, Weicheng, et al.
Veröffentlicht: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
von: Zhong, Xinrui, et al.
Veröffentlicht: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Xorbits: Automating Operator Tiling for Distributed Data Science
von: Lu, Weizheng, et al.
Veröffentlicht: (2023)
von: Lu, Weizheng, et al.
Veröffentlicht: (2023)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024) -
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025) -
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025) -
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
von: Xue, Chunyu, et al.
Veröffentlicht: (2026) -
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)