EmuGEMM (Benchmark Scripts): Fused Tensor Core Kernels for Precision Emulation in Matrix Multiplication
Fuente:
Zenodo
Guardado en:
| Autores principales: | Lu, Denghui, Maeder, Alexander, Ziogas, Alexandros Nikolaos |
|---|---|
| Formato: | Recurso digital |
| Publicado: |
Zenodo
2026
|
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quatrex SC25 AD: Simulation Inputs and Configuration Files
por: Ziogas, Alexandros Nikolaos, et al.
Publicado: (2025)
por: Ziogas, Alexandros Nikolaos, et al.
Publicado: (2025)
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
por: Kaniselvan, Manasa, et al.
Publicado: (2025)
por: Kaniselvan, Manasa, et al.
Publicado: (2025)
Electron-Electron Interactions in Device Simulation via Non-equilibrium Green's Functions and the GW Approximation
por: Deuschle, Leonard, et al.
Publicado: (2024)
por: Deuschle, Leonard, et al.
Publicado: (2024)
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
por: Yang, Shaoliang, et al.
Publicado: (2026)
por: Yang, Shaoliang, et al.
Publicado: (2026)
Learning the Electronic Hamiltonian of Large Atomic Structures
por: Xia, Chen Hao, et al.
Publicado: (2025)
por: Xia, Chen Hao, et al.
Publicado: (2025)
Single Neuromorphic Memristor closely Emulates Multiple Synaptic Mechanisms for Energy Efficient Neural Networks
por: Weilenmann, Christoph, et al.
Publicado: (2024)
por: Weilenmann, Christoph, et al.
Publicado: (2024)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
por: Li, Haomin, et al.
Publicado: (2026)
por: Li, Haomin, et al.
Publicado: (2026)
FlexEmu: Towards Flexible MCU Peripheral Emulation (Extended Version)
por: Lei, Chongqing, et al.
Publicado: (2025)
por: Lei, Chongqing, et al.
Publicado: (2025)
FV8: A Forced Execution JavaScript Engine for Detecting Evasive Techniques
por: Pantelaios, Nikolaos, et al.
Publicado: (2024)
por: Pantelaios, Nikolaos, et al.
Publicado: (2024)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
por: Gianinazzi, Lukas, et al.
Publicado: (2024)
por: Gianinazzi, Lukas, et al.
Publicado: (2024)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
por: Valpey, Benjamin, et al.
Publicado: (2025)
por: Valpey, Benjamin, et al.
Publicado: (2025)
Ab-initio Quantum Transport with the GW Approximation, 42,240 Atoms, and Sustained Exascale Performance
por: Vetsch, Nicolas, et al.
Publicado: (2025)
por: Vetsch, Nicolas, et al.
Publicado: (2025)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
por: Zhu, Honglin, et al.
Publicado: (2026)
por: Zhu, Honglin, et al.
Publicado: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
por: Da, Wei, et al.
Publicado: (2026)
por: Da, Wei, et al.
Publicado: (2026)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
Parallel Quadratic Selected Inversion in Quantum Transport Simulation
por: Maillou, Vincent, et al.
Publicado: (2026)
por: Maillou, Vincent, et al.
Publicado: (2026)
Cascading GEMM: High Precision from Low Precision
por: Parikh, Devangi N., et al.
Publicado: (2023)
por: Parikh, Devangi N., et al.
Publicado: (2023)
Acceleration of Atomistic NEGF: Algorithms, Parallelization, and Machine Learning
por: Luisier, Mathieu, et al.
Publicado: (2026)
por: Luisier, Mathieu, et al.
Publicado: (2026)
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
por: Uchino, Yuki, et al.
Publicado: (2026)
por: Uchino, Yuki, et al.
Publicado: (2026)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
por: Maillou, Vincent, et al.
Publicado: (2025)
por: Maillou, Vincent, et al.
Publicado: (2025)
NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation
por: Wan, Haoran, et al.
Publicado: (2026)
por: Wan, Haoran, et al.
Publicado: (2026)
EmuPlat: A Framework-Agnostic Platform for Quantum Hardware Emulation with Validated Transpiler-to-Pulse Pipeline
por: Ye, Jun, et al.
Publicado: (2025)
por: Ye, Jun, et al.
Publicado: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
por: Hu, Huanqi, et al.
Publicado: (2025)
por: Hu, Huanqi, et al.
Publicado: (2025)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
por: Nair, Harideep, et al.
Publicado: (2024)
por: Nair, Harideep, et al.
Publicado: (2024)
Fusing the Polyhedral and Tensor Compilers to Accelerate Scientific Computing Kernels
por: Qingzhi Liu, et al.
Publicado: (2025)
por: Qingzhi Liu, et al.
Publicado: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
por: Rout, Nikhil, et al.
Publicado: (2025)
por: Rout, Nikhil, et al.
Publicado: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
por: Xue, Weicheng, et al.
Publicado: (2025)
por: Xue, Weicheng, et al.
Publicado: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
por: Metere, Alfredo
Publicado: (2025)
por: Metere, Alfredo
Publicado: (2025)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
por: Remke, Stefan, et al.
Publicado: (2024)
por: Remke, Stefan, et al.
Publicado: (2024)
Emu: Generative Pretraining in Multimodality
por: Sun, Quan, et al.
Publicado: (2023)
por: Sun, Quan, et al.
Publicado: (2023)
CloudEmu: A Trace-Driven Cloud-Native Emulation Testbed for Vehicle Video Uplink over Cellular Networks
por: Torii, Takashi, et al.
Publicado: (2026)
por: Torii, Takashi, et al.
Publicado: (2026)
Manifold Gaussian Variational Bayes on the Precision Matrix
por: Magris, Martin, et al.
Publicado: (2022)
por: Magris, Martin, et al.
Publicado: (2022)
Fused3S: Fast Sparse Attention on Tensor Cores
por: Li, Zitong, et al.
Publicado: (2025)
por: Li, Zitong, et al.
Publicado: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
por: Zhao, Haisha, et al.
Publicado: (2025)
por: Zhao, Haisha, et al.
Publicado: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
por: Park, Gunho, et al.
Publicado: (2022)
por: Park, Gunho, et al.
Publicado: (2022)
FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
por: Zuo, Fei, et al.
Publicado: (2026)
por: Zuo, Fei, et al.
Publicado: (2026)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
por: Uchino, Yuki, et al.
Publicado: (2025)
por: Uchino, Yuki, et al.
Publicado: (2025)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
por: Shi, Jinliang, et al.
Publicado: (2024)
por: Shi, Jinliang, et al.
Publicado: (2024)
LP-GEMM: Integrating Layout Propagation into GEMM Operations
por: Carneiro, César Guedes, et al.
Publicado: (2026)
por: Carneiro, César Guedes, et al.
Publicado: (2026)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
por: Agarwal, Krish, et al.
Publicado: (2024)
por: Agarwal, Krish, et al.
Publicado: (2024)
Ejemplares similares
-
Quatrex SC25 AD: Simulation Inputs and Configuration Files
por: Ziogas, Alexandros Nikolaos, et al.
Publicado: (2025) -
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
por: Kaniselvan, Manasa, et al.
Publicado: (2025) -
Electron-Electron Interactions in Device Simulation via Non-equilibrium Green's Functions and the GW Approximation
por: Deuschle, Leonard, et al.
Publicado: (2024) -
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
por: Yang, Shaoliang, et al.
Publicado: (2026) -
Learning the Electronic Hamiltonian of Large Atomic Structures
por: Xia, Chen Hao, et al.
Publicado: (2025)