EmuGEMM (Benchmark Scripts): Fused Tensor Core Kernels for Precision Emulation in Matrix Multiplication
Fuente:
Zenodo
Saved in:
| Main Authors: | Lu, Denghui, Maeder, Alexander, Ziogas, Alexandros Nikolaos |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quatrex SC25 AD: Simulation Inputs and Configuration Files
by: Ziogas, Alexandros Nikolaos, et al.
Published: (2025)
by: Ziogas, Alexandros Nikolaos, et al.
Published: (2025)
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
by: Kaniselvan, Manasa, et al.
Published: (2025)
by: Kaniselvan, Manasa, et al.
Published: (2025)
Electron-Electron Interactions in Device Simulation via Non-equilibrium Green's Functions and the GW Approximation
by: Deuschle, Leonard, et al.
Published: (2024)
by: Deuschle, Leonard, et al.
Published: (2024)
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Learning the Electronic Hamiltonian of Large Atomic Structures
by: Xia, Chen Hao, et al.
Published: (2025)
by: Xia, Chen Hao, et al.
Published: (2025)
Single Neuromorphic Memristor closely Emulates Multiple Synaptic Mechanisms for Energy Efficient Neural Networks
by: Weilenmann, Christoph, et al.
Published: (2024)
by: Weilenmann, Christoph, et al.
Published: (2024)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
FlexEmu: Towards Flexible MCU Peripheral Emulation (Extended Version)
by: Lei, Chongqing, et al.
Published: (2025)
by: Lei, Chongqing, et al.
Published: (2025)
FV8: A Forced Execution JavaScript Engine for Detecting Evasive Techniques
by: Pantelaios, Nikolaos, et al.
Published: (2024)
by: Pantelaios, Nikolaos, et al.
Published: (2024)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
by: Gianinazzi, Lukas, et al.
Published: (2024)
by: Gianinazzi, Lukas, et al.
Published: (2024)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
by: Valpey, Benjamin, et al.
Published: (2025)
by: Valpey, Benjamin, et al.
Published: (2025)
Ab-initio Quantum Transport with the GW Approximation, 42,240 Atoms, and Sustained Exascale Performance
by: Vetsch, Nicolas, et al.
Published: (2025)
by: Vetsch, Nicolas, et al.
Published: (2025)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
by: Zhu, Honglin, et al.
Published: (2026)
by: Zhu, Honglin, et al.
Published: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
by: Da, Wei, et al.
Published: (2026)
by: Da, Wei, et al.
Published: (2026)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
Parallel Quadratic Selected Inversion in Quantum Transport Simulation
by: Maillou, Vincent, et al.
Published: (2026)
by: Maillou, Vincent, et al.
Published: (2026)
Cascading GEMM: High Precision from Low Precision
by: Parikh, Devangi N., et al.
Published: (2023)
by: Parikh, Devangi N., et al.
Published: (2023)
Acceleration of Atomistic NEGF: Algorithms, Parallelization, and Machine Learning
by: Luisier, Mathieu, et al.
Published: (2026)
by: Luisier, Mathieu, et al.
Published: (2026)
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
by: Maillou, Vincent, et al.
Published: (2025)
by: Maillou, Vincent, et al.
Published: (2025)
NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation
by: Wan, Haoran, et al.
Published: (2026)
by: Wan, Haoran, et al.
Published: (2026)
EmuPlat: A Framework-Agnostic Platform for Quantum Hardware Emulation with Validated Transpiler-to-Pulse Pipeline
by: Ye, Jun, et al.
Published: (2025)
by: Ye, Jun, et al.
Published: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Fusing the Polyhedral and Tensor Compilers to Accelerate Scientific Computing Kernels
by: Qingzhi Liu, et al.
Published: (2025)
by: Qingzhi Liu, et al.
Published: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
CloudEmu: A Trace-Driven Cloud-Native Emulation Testbed for Vehicle Video Uplink over Cellular Networks
by: Torii, Takashi, et al.
Published: (2026)
by: Torii, Takashi, et al.
Published: (2026)
Manifold Gaussian Variational Bayes on the Precision Matrix
by: Magris, Martin, et al.
Published: (2022)
by: Magris, Martin, et al.
Published: (2022)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)
by: Park, Gunho, et al.
Published: (2022)
FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
by: Zuo, Fei, et al.
Published: (2026)
by: Zuo, Fei, et al.
Published: (2026)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
by: Shi, Jinliang, et al.
Published: (2024)
by: Shi, Jinliang, et al.
Published: (2024)
LP-GEMM: Integrating Layout Propagation into GEMM Operations
by: Carneiro, César Guedes, et al.
Published: (2026)
by: Carneiro, César Guedes, et al.
Published: (2026)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
by: Agarwal, Krish, et al.
Published: (2024)
by: Agarwal, Krish, et al.
Published: (2024)
Similar Items
-
Quatrex SC25 AD: Simulation Inputs and Configuration Files
by: Ziogas, Alexandros Nikolaos, et al.
Published: (2025) -
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
by: Kaniselvan, Manasa, et al.
Published: (2025) -
Electron-Electron Interactions in Device Simulation via Non-equilibrium Green's Functions and the GW Approximation
by: Deuschle, Leonard, et al.
Published: (2024) -
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
by: Yang, Shaoliang, et al.
Published: (2026) -
Learning the Electronic Hamiltonian of Large Atomic Structures
by: Xia, Chen Hao, et al.
Published: (2025)