Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Peichen, Xu, Shuotao, Wang, Yang, Yang, Fan, Yang, Mao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hawkeye: Reproducing GPU-Level Non-Determinism
von: Badash, Erez, et al.
Veröffentlicht: (2026)
von: Badash, Erez, et al.
Veröffentlicht: (2026)
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
An Open-Source Framework for Efficient Numerically-Tailored Computations
von: Ledoux, Louis, et al.
Veröffentlicht: (2024)
von: Ledoux, Louis, et al.
Veröffentlicht: (2024)
Accurate Models of NVIDIA Tensor Cores
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
Accurate Block Quantization in LLMs with Outliers
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
von: Trukhanov, Nikita, et al.
Veröffentlicht: (2024)
Design and accuracy trade-offs in Computational Statistics
von: Xu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Xu, Tiancheng, et al.
Veröffentlicht: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
Efficient FRW Transitions via Stochastic Finite Differences for Handling Non-Stratified Dielectrics
von: Huang, Jiechen, et al.
Veröffentlicht: (2025)
von: Huang, Jiechen, et al.
Veröffentlicht: (2025)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
von: Li, Guoyu, et al.
Veröffentlicht: (2025)
von: Li, Guoyu, et al.
Veröffentlicht: (2025)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
von: Natesh, Vikas, et al.
Veröffentlicht: (2025)
Efficient FIR filtering with Bit Layer Multiply Accumulator
von: Liguori, Vincenzo
Veröffentlicht: (2024)
von: Liguori, Vincenzo
Veröffentlicht: (2024)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
von: Shao, Kunming, et al.
Veröffentlicht: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
Leveraging Highly Approximated Multipliers in DNN Inference
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
von: Xie, Xilong, et al.
Veröffentlicht: (2025)
Approximate Multiplier Induced Error Propagation in Deep Neural Networks
von: Alahakoon, A. M. H. H., et al.
Veröffentlicht: (2025)
von: Alahakoon, A. M. H. H., et al.
Veröffentlicht: (2025)
Bit-Flip Fault Attack: Crushing Graph Neural Networks via Gradual Bit Search
von: Abharian, Sanaz Kazemi, et al.
Veröffentlicht: (2025)
von: Abharian, Sanaz Kazemi, et al.
Veröffentlicht: (2025)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
von: Shende, Omkar B, et al.
Veröffentlicht: (2026)
von: Shende, Omkar B, et al.
Veröffentlicht: (2026)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
MATLAB Simulator of Level-Index Arithmetic
von: Mikaitis, Mantas
Veröffentlicht: (2024)
von: Mikaitis, Mantas
Veröffentlicht: (2024)
Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
von: Croci, M., et al.
Veröffentlicht: (2024)
von: Croci, M., et al.
Veröffentlicht: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
von: Zhang, Tao, et al.
Veröffentlicht: (2026)
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
von: Kim, Jiin, et al.
Veröffentlicht: (2025)
von: Kim, Jiin, et al.
Veröffentlicht: (2025)
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
von: Pan, Yunjie, et al.
Veröffentlicht: (2026)
von: Pan, Yunjie, et al.
Veröffentlicht: (2026)
RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space Reduction
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
Bespoke Approximation of Multiplication-Accumulation and Activation Targeting Printed Multilayer Perceptrons
von: Afentaki, Florentia, et al.
Veröffentlicht: (2023)
von: Afentaki, Florentia, et al.
Veröffentlicht: (2023)
Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
von: Wu, Qizhe, et al.
Veröffentlicht: (2025)
von: Wu, Qizhe, et al.
Veröffentlicht: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
MORCIC: Model Order Reduction Techniques for Electromagnetic Models of Integrated Circuits
von: Garyfallou, Dimitrios, et al.
Veröffentlicht: (2023)
von: Garyfallou, Dimitrios, et al.
Veröffentlicht: (2023)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
von: Pelke, Rebecca, et al.
Veröffentlicht: (2025)
von: Pelke, Rebecca, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hawkeye: Reproducing GPU-Level Non-Determinism
von: Badash, Erez, et al.
Veröffentlicht: (2026) -
eXmY: A Data Type and Technique for Arbitrary Bit Precision Quantization
von: Agrawal, Aditya, et al.
Veröffentlicht: (2024) -
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
von: Xue, Chenhao, et al.
Veröffentlicht: (2025) -
An Open-Source Framework for Efficient Numerically-Tailored Computations
von: Ledoux, Louis, et al.
Veröffentlicht: (2024) -
Accurate Models of NVIDIA Tensor Cores
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)