Accurate Models of NVIDIA Tensor Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khattak, Faizan A., Mikaitis, Mantas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MATLAB Simulator of Level-Index Arithmetic
von: Mikaitis, Mantas
Veröffentlicht: (2024)
von: Mikaitis, Mantas
Veröffentlicht: (2024)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
von: Croci, M., et al.
Veröffentlicht: (2024)
von: Croci, M., et al.
Veröffentlicht: (2024)
An Open-Source Framework for Efficient Numerically-Tailored Computations
von: Ledoux, Louis, et al.
Veröffentlicht: (2024)
von: Ledoux, Louis, et al.
Veröffentlicht: (2024)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
von: Mukunoki, Daichi
Veröffentlicht: (2025)
von: Mukunoki, Daichi
Veröffentlicht: (2025)
Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
Inexactness and Correction of Floating-Point Reciprocal, Division and Square Root
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
von: Dutton, Lucas M., et al.
Veröffentlicht: (2024)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
Accurate Block Quantization in LLMs with Outliers
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
HYLU: Hybrid Parallel Sparse LU Factorization
von: Chen, Xiaoming
Veröffentlicht: (2025)
von: Chen, Xiaoming
Veröffentlicht: (2025)
Design and accuracy trade-offs in Computational Statistics
von: Xu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Xu, Tiancheng, et al.
Veröffentlicht: (2025)
Efficient FRW Transitions via Stochastic Finite Differences for Handling Non-Stratified Dielectrics
von: Huang, Jiechen, et al.
Veröffentlicht: (2025)
von: Huang, Jiechen, et al.
Veröffentlicht: (2025)
A Hybrid Residue Floating Numerical Architecture for High Precision Arithmetic on FPGAs
von: Darvishi, Mostafa
Veröffentlicht: (2025)
von: Darvishi, Mostafa
Veröffentlicht: (2025)
Analysis of Floating-Point Matrix Multiplication Computed via Integer Arithmetic
von: Abdelfattah, Ahmad, et al.
Veröffentlicht: (2025)
von: Abdelfattah, Ahmad, et al.
Veröffentlicht: (2025)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
von: Ltaief, Hatem, et al.
Veröffentlicht: (2024)
Evaluation of POSIT Arithmetic with Accelerators
von: Nakasato, Naohito, et al.
Veröffentlicht: (2024)
von: Nakasato, Naohito, et al.
Veröffentlicht: (2024)
Fast and energy-efficient derivatives risk analysis: Streaming option Greeks on Xilinx and Intel FPGAs
von: Klaisoongnoen, Mark, et al.
Veröffentlicht: (2022)
von: Klaisoongnoen, Mark, et al.
Veröffentlicht: (2022)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
von: Valpey, Benjamin, et al.
Veröffentlicht: (2025)
von: Valpey, Benjamin, et al.
Veröffentlicht: (2025)
MORCIC: Model Order Reduction Techniques for Electromagnetic Models of Integrated Circuits
von: Garyfallou, Dimitrios, et al.
Veröffentlicht: (2023)
von: Garyfallou, Dimitrios, et al.
Veröffentlicht: (2023)
Accuracy of Mathematical Functions in Julia
von: Mikaitis, Mantas, et al.
Veröffentlicht: (2025)
von: Mikaitis, Mantas, et al.
Veröffentlicht: (2025)
LeGend: A Data-Driven Framework for Lemma Generation in Hardware Model Checking
von: Miao, Mingkai, et al.
Veröffentlicht: (2026)
von: Miao, Mingkai, et al.
Veröffentlicht: (2026)
C2HLSC: Leveraging Large Language Models to Bridge the Software-to-Hardware Design Gap
von: Collini, Luca, et al.
Veröffentlicht: (2024)
von: Collini, Luca, et al.
Veröffentlicht: (2024)
Hawkeye: Reproducing GPU-Level Non-Determinism
von: Badash, Erez, et al.
Veröffentlicht: (2026)
von: Badash, Erez, et al.
Veröffentlicht: (2026)
A low-rank balanced truncation approach for large-scale RLCk model order reduction based on extended Krylov subspace and a frequency-aware convergence criterion
von: Giamouzis, Christos, et al.
Veröffentlicht: (2024)
von: Giamouzis, Christos, et al.
Veröffentlicht: (2024)
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
FLAG: Formal and LLM-assisted SVA Generation for Formal Specifications of On-Chip Communication Protocols
von: Shih, Yu-An, et al.
Veröffentlicht: (2025)
von: Shih, Yu-An, et al.
Veröffentlicht: (2025)
Offloading Data Center Tax
von: Revankar, Akshay, et al.
Veröffentlicht: (2025)
von: Revankar, Akshay, et al.
Veröffentlicht: (2025)
A Vertically Integrated Framework for Templatized Chip Design
von: Kim, Jeongeun, et al.
Veröffentlicht: (2025)
von: Kim, Jeongeun, et al.
Veröffentlicht: (2025)
Structural Mutation Based Differential Testing for FPGA Logic Synthesis Compilers
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
Scalable Software Testing in Fast Virtual Platforms: Leveraging SystemC, QEMU and Containerization
von: Jünger, Lukas, et al.
Veröffentlicht: (2025)
von: Jünger, Lukas, et al.
Veröffentlicht: (2025)
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
von: Vavelidou, Ioanna, et al.
Veröffentlicht: (2026)
von: Vavelidou, Ioanna, et al.
Veröffentlicht: (2026)
Using LLMs to Facilitate Formal Verification of RTL
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
EquivFusion: Unifying Hardware Equivalence Checking from Algorithms to Netlists via MLIR
von: Zhu, Jiaying, et al.
Veröffentlicht: (2026)
von: Zhu, Jiaying, et al.
Veröffentlicht: (2026)
UVMarvel: an Automated LLM-aided UVM Machine for Subsystem-level RTL Verification
von: Ye, Junhao, et al.
Veröffentlicht: (2026)
von: Ye, Junhao, et al.
Veröffentlicht: (2026)
MEIC: Re-thinking RTL Debug Automation using LLMs
von: Xu, Ke, et al.
Veröffentlicht: (2024)
von: Xu, Ke, et al.
Veröffentlicht: (2024)
AutoINV: Automated Invariant Generation Framework for Formal Verification on High-Level Synthesis Designs
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
Quantifying Uncertainty in FMEDA Safety Metrics: An Error Propagation Approach for Enhanced ASIC Verification
von: Armato, Antonino, et al.
Veröffentlicht: (2026)
von: Armato, Antonino, et al.
Veröffentlicht: (2026)
FormalRTL: Verified RTL Synthesis at Scale
von: Li, Kezhi, et al.
Veröffentlicht: (2026)
von: Li, Kezhi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MATLAB Simulator of Level-Index Arithmetic
von: Mikaitis, Mantas
Veröffentlicht: (2024) -
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025) -
Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
von: Croci, M., et al.
Veröffentlicht: (2024) -
An Open-Source Framework for Efficient Numerically-Tailored Computations
von: Ledoux, Louis, et al.
Veröffentlicht: (2024) -
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
von: Mukunoki, Daichi
Veröffentlicht: (2025)