HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Krish, Astra, Rishi, Hoque, Adnan, Srivatsa, Mudhakar, Ganti, Raghu, Wright, Less, Chen, Sijia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
TP-Aware Dequantization
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)
by: Gu, Qiqi, et al.
Published: (2026)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
PICO: Accelerating All k-Core Paradigms on GPU
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025)
by: GU, Qiqi, et al.
Published: (2025)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
by: Huang, Ziyu, et al.
Published: (2025)
by: Huang, Ziyu, et al.
Published: (2025)
Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search
by: Meneses, Enzo, et al.
Published: (2026)
by: Meneses, Enzo, et al.
Published: (2026)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
by: Schwarz, Angelika, et al.
Published: (2025)
by: Schwarz, Angelika, et al.
Published: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
by: Liu, Zhibang, et al.
Published: (2025)
by: Liu, Zhibang, et al.
Published: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026)
by: Pacheco, Daniel, et al.
Published: (2026)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
by: Epstein, Elliot L., et al.
Published: (2026)
by: Epstein, Elliot L., et al.
Published: (2026)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
by: Ottaviano, Alessandro, et al.
Published: (2025)
by: Ottaviano, Alessandro, et al.
Published: (2025)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
by: Carrasco, Roberto, et al.
Published: (2026)
by: Carrasco, Roberto, et al.
Published: (2026)
BLEST: Blazingly Efficient BFS using Tensor Cores
by: Elbek, Deniz, et al.
Published: (2025)
by: Elbek, Deniz, et al.
Published: (2025)
Experimental Evaluation of Distributed k-Core Decomposition
by: Guo, Bin, et al.
Published: (2024)
by: Guo, Bin, et al.
Published: (2024)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Decentralized and Self-adaptive Core Maintenance on Temporal Graphs
by: Rucci, Davide, et al.
Published: (2025)
by: Rucci, Davide, et al.
Published: (2025)
Parallel Order-Based Core Maintenance in Dynamic Graphs
by: Guo, Bin, et al.
Published: (2022)
by: Guo, Bin, et al.
Published: (2022)
A Parallel Scan Algorithm in the Tensor Core Unit Model
by: Zouzias, Anastasios, et al.
Published: (2024)
by: Zouzias, Anastasios, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Mean field optimal Core Allocation across Malleable jobs
by: Li, Zhouzi, et al.
Published: (2026)
by: Li, Zhouzi, et al.
Published: (2026)
Exploring Fine-grained Task Parallelism on Simultaneous Multithreading Cores
by: Los, Denis, et al.
Published: (2024)
by: Los, Denis, et al.
Published: (2024)
Similar Items
-
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024) -
TP-Aware Dequantization
by: Hoque, Adnan, et al.
Published: (2024) -
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025) -
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024) -
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)