FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xinyi, Li, Ang, Fang, Bo, Swirydowicz, Katarzyna, Laguna, Ignacio, Gopalakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Accelerating CRONet on AMD Versal AIE-ML Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2026)
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2026)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
AMD Versal Implementations of FAM and SSCA Estimators
von: Li, Carol Jingyi, et al.
Veröffentlicht: (2025)
von: Li, Carol Jingyi, et al.
Veröffentlicht: (2025)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2025)
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2025)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
von: Nagendra, Savinay
Veröffentlicht: (2024)
von: Nagendra, Savinay
Veröffentlicht: (2024)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
von: Valpey, Benjamin, et al.
Veröffentlicht: (2025)
von: Valpey, Benjamin, et al.
Veröffentlicht: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
NOVA: Coordinated Test Selection and Bayes-Optimized Constrained Randomization for Accelerated Coverage Closure
von: Peng, Weijie, et al.
Veröffentlicht: (2025)
von: Peng, Weijie, et al.
Veröffentlicht: (2025)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
von: Wang, Kainan, et al.
Veröffentlicht: (2025)
von: Wang, Kainan, et al.
Veröffentlicht: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
von: Antunes, Pedro, et al.
Veröffentlicht: (2026)
von: Antunes, Pedro, et al.
Veröffentlicht: (2026)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A, et al.
Veröffentlicht: (2025)
Duet: Creating Harmony between Processors and Embedded FPGAs
von: Li, Ang, et al.
Veröffentlicht: (2023)
von: Li, Ang, et al.
Veröffentlicht: (2023)
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design
von: You, Haoran, et al.
Veröffentlicht: (2021)
von: You, Haoran, et al.
Veröffentlicht: (2021)
Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference
von: Xia, Tianhua, et al.
Veröffentlicht: (2023)
von: Xia, Tianhua, et al.
Veröffentlicht: (2023)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
von: Del Sozzo, Emanuele, et al.
Veröffentlicht: (2026)
von: Del Sozzo, Emanuele, et al.
Veröffentlicht: (2026)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
von: Rösti, André, et al.
Veröffentlicht: (2025)
von: Rösti, André, et al.
Veröffentlicht: (2025)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
von: Le, Tran Xuan Hieu, et al.
Veröffentlicht: (2024)
von: Le, Tran Xuan Hieu, et al.
Veröffentlicht: (2024)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Real-time Object Detection and Associated Hardware Accelerators Targeting Autonomous Vehicles: A Review
von: Sali, Safa, et al.
Veröffentlicht: (2025)
von: Sali, Safa, et al.
Veröffentlicht: (2025)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
von: Li, Jindong, et al.
Veröffentlicht: (2024)
von: Li, Jindong, et al.
Veröffentlicht: (2024)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
von: Makin, Yashasvi, et al.
Veröffentlicht: (2025)
von: Makin, Yashasvi, et al.
Veröffentlicht: (2025)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
von: Chen, Bo-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Bo-Yu, et al.
Veröffentlicht: (2025)
Accurate Models of NVIDIA Tensor Cores
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
von: Khattak, Faizan A., et al.
Veröffentlicht: (2025)
NeuroFlex: Column-Exact ANN-SNN Co-Execution Accelerator with Cost-Guided Scheduling
von: Manjunath, Varun, et al.
Veröffentlicht: (2025)
von: Manjunath, Varun, et al.
Veröffentlicht: (2025)
Hybrid Photonic-digital Accelerator for Attention Mechanism
von: Li, Huize, et al.
Veröffentlicht: (2025)
von: Li, Huize, et al.
Veröffentlicht: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023) -
Accelerating CRONet on AMD Versal AIE-ML Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2026) -
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025) -
AMD Versal Implementations of FAM and SSCA Estimators
von: Li, Carol Jingyi, et al.
Veröffentlicht: (2025) -
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2025)