GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
Fuente:
arXiv
Guardado en:
| Autores principales: | Mhatre, Kaustubh, Taka, Endri, Arora, Aman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Accelerating CRONet on AMD Versal AIE-ML Engines
por: Mhatre, Kaustubh, et al.
Publicado: (2026)
por: Mhatre, Kaustubh, et al.
Publicado: (2026)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
por: Taka, Endri, et al.
Publicado: (2024)
por: Taka, Endri, et al.
Publicado: (2024)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
por: Taka, Endri, et al.
Publicado: (2025)
por: Taka, Endri, et al.
Publicado: (2025)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
por: Papalamprou, Ilias, et al.
Publicado: (2025)
por: Papalamprou, Ilias, et al.
Publicado: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
por: Taka, Endri, et al.
Publicado: (2025)
por: Taka, Endri, et al.
Publicado: (2025)
AMD Versal Implementations of FAM and SSCA Estimators
por: Li, Carol Jingyi, et al.
Publicado: (2025)
por: Li, Carol Jingyi, et al.
Publicado: (2025)
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
por: Ohno, Ayumi, et al.
Publicado: (2025)
por: Ohno, Ayumi, et al.
Publicado: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
por: Shimamura, Kotaro, et al.
Publicado: (2025)
por: Shimamura, Kotaro, et al.
Publicado: (2025)
CAT: Customized Transformer Accelerator Framework on Versal ACAP
por: Zhang, Wenbo, et al.
Publicado: (2024)
por: Zhang, Wenbo, et al.
Publicado: (2024)
Enabling Mixed criticality applications for the Versal AI-Engines
por: Sprave, Vincent, et al.
Publicado: (2026)
por: Sprave, Vincent, et al.
Publicado: (2026)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
por: Guo, Cong, et al.
Publicado: (2025)
por: Guo, Cong, et al.
Publicado: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
por: Li, Haomin, et al.
Publicado: (2026)
por: Li, Haomin, et al.
Publicado: (2026)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
por: Karami, Rachid, et al.
Publicado: (2024)
por: Karami, Rachid, et al.
Publicado: (2024)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
por: Ma, Siyuan, et al.
Publicado: (2025)
por: Ma, Siyuan, et al.
Publicado: (2025)
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
por: Sui, Bingcai, et al.
Publicado: (2024)
por: Sui, Bingcai, et al.
Publicado: (2024)
CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems
por: Wu, Qihang, et al.
Publicado: (2026)
por: Wu, Qihang, et al.
Publicado: (2026)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
por: Grailoo, M., et al.
Publicado: (2026)
por: Grailoo, M., et al.
Publicado: (2026)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
por: Dai, Tuo, et al.
Publicado: (2024)
por: Dai, Tuo, et al.
Publicado: (2024)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
por: Rösti, André, et al.
Publicado: (2025)
por: Rösti, André, et al.
Publicado: (2025)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
por: Li, Xinyi, et al.
Publicado: (2024)
por: Li, Xinyi, et al.
Publicado: (2024)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
por: Boutros, Andrew, et al.
Publicado: (2024)
por: Boutros, Andrew, et al.
Publicado: (2024)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
por: Patel, Vihaan, et al.
Publicado: (2026)
por: Patel, Vihaan, et al.
Publicado: (2026)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
por: Wang, Erwei, et al.
Publicado: (2025)
por: Wang, Erwei, et al.
Publicado: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
por: Cammarata, Danilo, et al.
Publicado: (2026)
por: Cammarata, Danilo, et al.
Publicado: (2026)
AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
por: Danopoulos, Dimitrios, et al.
Publicado: (2025)
por: Danopoulos, Dimitrios, et al.
Publicado: (2025)
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
Towards Employing FPGA and ASIP Acceleration to Enable Onboard AI/ML in Space Applications
por: Leon, Vasileios, et al.
Publicado: (2025)
por: Leon, Vasileios, et al.
Publicado: (2025)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
por: Wu, Qizhe, et al.
Publicado: (2024)
por: Wu, Qizhe, et al.
Publicado: (2024)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
por: Petropoulos, Anastasios, et al.
Publicado: (2025)
por: Petropoulos, Anastasios, et al.
Publicado: (2025)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
por: Ta, Tuan, et al.
Publicado: (2025)
por: Ta, Tuan, et al.
Publicado: (2025)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
por: Li, Guoyu, et al.
Publicado: (2025)
por: Li, Guoyu, et al.
Publicado: (2025)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
por: Li, Tenglong, et al.
Publicado: (2025)
por: Li, Tenglong, et al.
Publicado: (2025)
Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators
por: Yik, Jason, et al.
Publicado: (2025)
por: Yik, Jason, et al.
Publicado: (2025)
ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
por: Jain, Devansh, et al.
Publicado: (2025)
por: Jain, Devansh, et al.
Publicado: (2025)
Record Acceleration of the Two-Dimensional Ising Model Using High-Performance Wafer Scale Engine
por: Van Essendelft, Dirk, et al.
Publicado: (2024)
por: Van Essendelft, Dirk, et al.
Publicado: (2024)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
por: Nair, Harideep, et al.
Publicado: (2024)
por: Nair, Harideep, et al.
Publicado: (2024)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
por: Hu, Jiajun, et al.
Publicado: (2025)
por: Hu, Jiajun, et al.
Publicado: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
Ejemplares similares
-
Accelerating CRONet on AMD Versal AIE-ML Engines
por: Mhatre, Kaustubh, et al.
Publicado: (2026) -
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
por: Taka, Endri, et al.
Publicado: (2024) -
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
por: Taka, Endri, et al.
Publicado: (2025) -
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
por: Papalamprou, Ilias, et al.
Publicado: (2025) -
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
por: Taka, Endri, et al.
Publicado: (2025)