AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
Fuente:
arXiv
Saved in:
| Main Author: | Stankovic, Aleksandar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
by: Belias, François, et al.
Published: (2025)
by: Belias, François, et al.
Published: (2025)
FuseSampleAgg: Fused Neighbor Sampling and Aggregation for Mini-batch GNNs
by: Stanković, Aleksandar
Published: (2025)
by: Stanković, Aleksandar
Published: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
by: Baronio, Carlo, et al.
Published: (2025)
by: Baronio, Carlo, et al.
Published: (2025)
SENSEi: Input-Sensitive Compilation for Accelerating GNNs
by: Lenadora, Damitha, et al.
Published: (2023)
by: Lenadora, Damitha, et al.
Published: (2023)
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
by: Xiang, Lizhi, et al.
Published: (2025)
by: Xiang, Lizhi, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
by: Yao, Feiyu, et al.
Published: (2026)
by: Yao, Feiyu, et al.
Published: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
by: Lipshitz, Baraq, et al.
Published: (2025)
by: Lipshitz, Baraq, et al.
Published: (2025)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025)
by: You, Bozhi, et al.
Published: (2025)
P-MOSS: Scheduling Main-Memory Indexes Over NUMA Servers Using Next Token Prediction
by: Rayhan, Yeasir, et al.
Published: (2024)
by: Rayhan, Yeasir, et al.
Published: (2024)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
by: Duan, Shukai, et al.
Published: (2024)
by: Duan, Shukai, et al.
Published: (2024)
EARL: Energy-Aware Optimization of Liquid State Machines for Pervasive AI
by: Iqbal, Zain, et al.
Published: (2026)
by: Iqbal, Zain, et al.
Published: (2026)
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
by: Ullah, Irfan, et al.
Published: (2026)
by: Ullah, Irfan, et al.
Published: (2026)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
by: Song, Yingchen, et al.
Published: (2025)
by: Song, Yingchen, et al.
Published: (2025)
Efficient Solving of Large Single Input Superstate Decomposable Markovian Decision Process
by: Mahjoub, Youssef Ait El, et al.
Published: (2025)
by: Mahjoub, Youssef Ait El, et al.
Published: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
by: Xue, Leyang, et al.
Published: (2024)
by: Xue, Leyang, et al.
Published: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024)
by: Sarkar, Aishwarya, et al.
Published: (2024)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
by: Wu, Dan, et al.
Published: (2023)
by: Wu, Dan, et al.
Published: (2023)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
by: Liu, Jiashuo, et al.
Published: (2025)
by: Liu, Jiashuo, et al.
Published: (2025)
Toward Smart Scheduling in Tapis
by: Stubbs, Joe, et al.
Published: (2024)
by: Stubbs, Joe, et al.
Published: (2024)
Hybrid SpMM SC26 AD
by: Wang, Haoyu
Published: (2026)
by: Wang, Haoyu
Published: (2026)
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
by: Tomczak, Nathaniel, et al.
Published: (2025)
by: Tomczak, Nathaniel, et al.
Published: (2025)
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
by: Zhao, Mark, et al.
Published: (2024)
by: Zhao, Mark, et al.
Published: (2024)
Similar Items
-
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2024) -
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
by: Liu, Jiahui, et al.
Published: (2024) -
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025) -
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
by: Belias, François, et al.
Published: (2025) -
FuseSampleAgg: Fused Neighbor Sampling and Aggregation for Mini-batch GNNs
by: Stanković, Aleksandar
Published: (2025)