DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Mukunoki, Daichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)
Accurate Models of NVIDIA Tensor Cores
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
GreenMalloc: Allocator Optimisation for Industrial Workloads
von: Dakhama, Aidan, et al.
Veröffentlicht: (2025)
von: Dakhama, Aidan, et al.
Veröffentlicht: (2025)
Bridging the Gap: Physical PCI Device Integration Into SystemC-TLM Virtual Platforms
von: Bosbach, Nils, et al.
Veröffentlicht: (2025)
von: Bosbach, Nils, et al.
Veröffentlicht: (2025)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2024)
von: Esfahani, Mohsen Koohi, et al.
Veröffentlicht: (2024)
MATLAB Simulator of Level-Index Arithmetic
von: Mikaitis, Mantas
Veröffentlicht: (2024)
von: Mikaitis, Mantas
Veröffentlicht: (2024)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
A Hybrid Residue Floating Numerical Architecture for High Precision Arithmetic on FPGAs
von: Darvishi, Mostafa
Veröffentlicht: (2025)
von: Darvishi, Mostafa
Veröffentlicht: (2025)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
von: Perdomo, Elias, et al.
Veröffentlicht: (2024)
von: Perdomo, Elias, et al.
Veröffentlicht: (2024)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
von: Su, Zhongling, et al.
Veröffentlicht: (2025)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
Inexactness and Correction of Floating-Point Reciprocal, Division and Square Root
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
Evaluation of POSIT Arithmetic with Accelerators
von: Nakasato, Naohito, et al.
Veröffentlicht: (2024)
von: Nakasato, Naohito, et al.
Veröffentlicht: (2024)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
von: Liu, Songze, et al.
Veröffentlicht: (2025)
von: Liu, Songze, et al.
Veröffentlicht: (2025)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
von: Croci, M., et al.
Veröffentlicht: (2024)
von: Croci, M., et al.
Veröffentlicht: (2024)
Heterogeneous Memory Benchmarking Toolkit
von: Ghaemi, Golsana, et al.
Veröffentlicht: (2025)
von: Ghaemi, Golsana, et al.
Veröffentlicht: (2025)
Recurrent CircuitSAT Sampling for Sequential Circuits
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
Introducing the Arm-membench Throughput Benchmark
von: Burth, Cyrill, et al.
Veröffentlicht: (2025)
von: Burth, Cyrill, et al.
Veröffentlicht: (2025)
Enhancing software-hardware co-design for HEP by low-overhead profiling of single- and multi-threaded programs on diverse architectures with Adaptyst
von: Graczyk, Maksymilian, et al.
Veröffentlicht: (2025)
von: Graczyk, Maksymilian, et al.
Veröffentlicht: (2025)
AI Load Dynamics--A Power Electronics Perspective
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
von: Qararyah, Fareed, et al.
Veröffentlicht: (2025)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
SCALE-Sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026) -
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024) -
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
von: Zou, Jiaxiang, et al.
Veröffentlicht: (2026) -
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
Faster Inference of LLMs using FP8 on the Intel Gaudi
von: Lee, Joonhyung, et al.
Veröffentlicht: (2025)