Accelerating CRONet on AMD Versal AIE-ML Engines
Fuente:
arXiv
Guardado en:
| Autores principales: | Mhatre, Kaustubh, Tewari, Vedant, Ray, Aditya, Khan, Farhan, Olabiyi, Ridwan, Iquebal, Ashif, Arora, Aman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
por: Mhatre, Kaustubh, et al.
Publicado: (2025)
por: Mhatre, Kaustubh, et al.
Publicado: (2025)
AMD Versal Implementations of FAM and SSCA Estimators
por: Li, Carol Jingyi, et al.
Publicado: (2025)
por: Li, Carol Jingyi, et al.
Publicado: (2025)
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
por: Ohno, Ayumi, et al.
Publicado: (2025)
por: Ohno, Ayumi, et al.
Publicado: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
por: Shimamura, Kotaro, et al.
Publicado: (2025)
por: Shimamura, Kotaro, et al.
Publicado: (2025)
CAT: Customized Transformer Accelerator Framework on Versal ACAP
por: Zhang, Wenbo, et al.
Publicado: (2024)
por: Zhang, Wenbo, et al.
Publicado: (2024)
AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
por: Danopoulos, Dimitrios, et al.
Publicado: (2025)
por: Danopoulos, Dimitrios, et al.
Publicado: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
por: Ma, Siyuan, et al.
Publicado: (2025)
por: Ma, Siyuan, et al.
Publicado: (2025)
Enabling Mixed criticality applications for the Versal AI-Engines
por: Sprave, Vincent, et al.
Publicado: (2026)
por: Sprave, Vincent, et al.
Publicado: (2026)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
por: Taka, Endri, et al.
Publicado: (2024)
por: Taka, Endri, et al.
Publicado: (2024)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
por: Papalamprou, Ilias, et al.
Publicado: (2025)
por: Papalamprou, Ilias, et al.
Publicado: (2025)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
por: Rösti, André, et al.
Publicado: (2025)
por: Rösti, André, et al.
Publicado: (2025)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
por: Li, Xinyi, et al.
Publicado: (2024)
por: Li, Xinyi, et al.
Publicado: (2024)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
por: Patel, Vihaan, et al.
Publicado: (2026)
por: Patel, Vihaan, et al.
Publicado: (2026)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
por: Boutros, Andrew, et al.
Publicado: (2024)
por: Boutros, Andrew, et al.
Publicado: (2024)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
por: Taka, Endri, et al.
Publicado: (2025)
por: Taka, Endri, et al.
Publicado: (2025)
CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems
por: Wu, Qihang, et al.
Publicado: (2026)
por: Wu, Qihang, et al.
Publicado: (2026)
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2023)
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
por: Dai, Tuo, et al.
Publicado: (2024)
por: Dai, Tuo, et al.
Publicado: (2024)
ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
por: Jain, Devansh, et al.
Publicado: (2025)
por: Jain, Devansh, et al.
Publicado: (2025)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
por: Hu, Jiajun, et al.
Publicado: (2025)
por: Hu, Jiajun, et al.
Publicado: (2025)
TYTAN: Taylor-series based Non-Linear Activation Engine for Deep Learning Accelerators
por: Pramanik, Soham, et al.
Publicado: (2025)
por: Pramanik, Soham, et al.
Publicado: (2025)
LogicSparse: Enabling Engine-Free Unstructured Sparsity for Quantised Deep-learning Accelerators
por: Li, Changhong, et al.
Publicado: (2025)
por: Li, Changhong, et al.
Publicado: (2025)
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
por: Morales, José Juan Hernández, et al.
Publicado: (2026)
por: Morales, José Juan Hernández, et al.
Publicado: (2026)
From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators
por: Xie, Jiahan, et al.
Publicado: (2025)
por: Xie, Jiahan, et al.
Publicado: (2025)
Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs
por: Lowe, Sean, et al.
Publicado: (2026)
por: Lowe, Sean, et al.
Publicado: (2026)
Towards Employing FPGA and ASIP Acceleration to Enable Onboard AI/ML in Space Applications
por: Leon, Vasileios, et al.
Publicado: (2025)
por: Leon, Vasileios, et al.
Publicado: (2025)
EA4RCA:Efficient AIE accelerator design framework for Regular Communication-Avoiding Algorithm
por: Zhang, W. B., et al.
Publicado: (2024)
por: Zhang, W. B., et al.
Publicado: (2024)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
por: Li, Tenglong, et al.
Publicado: (2025)
por: Li, Tenglong, et al.
Publicado: (2025)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
por: Li, Guoyu, et al.
Publicado: (2025)
por: Li, Guoyu, et al.
Publicado: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
por: Yi, Xiaoling, et al.
Publicado: (2025)
por: Yi, Xiaoling, et al.
Publicado: (2025)
Accelerating Detailed Routing Convergence through Offline Reinforcement Learning
por: Khan, Afsara, et al.
Publicado: (2025)
por: Khan, Afsara, et al.
Publicado: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
por: Qararyah, Fareed, et al.
Publicado: (2025)
por: Qararyah, Fareed, et al.
Publicado: (2025)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2026)
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2026)
Discovering Governing Equations in the Presence of Uncertainty
por: Olabiyi, Ridwan, et al.
Publicado: (2025)
por: Olabiyi, Ridwan, et al.
Publicado: (2025)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
por: Li, Enlai, et al.
Publicado: (2026)
por: Li, Enlai, et al.
Publicado: (2026)
CarbonPATH: Carbon-aware pathfinding and architecture optimization for chiplet-based AI systems
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2026)
por: Sudarshan, Chetan Choppali, et al.
Publicado: (2026)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
por: Peng, Hongwu, et al.
Publicado: (2023)
por: Peng, Hongwu, et al.
Publicado: (2023)
Accelerating Time Series Analysis via Processing using Non-Volatile Memories
por: Fernandez, Ivan, et al.
Publicado: (2022)
por: Fernandez, Ivan, et al.
Publicado: (2022)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
por: Agrawal, Anirudha, et al.
Publicado: (2024)
por: Agrawal, Anirudha, et al.
Publicado: (2024)
Ejemplares similares
-
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
por: Mhatre, Kaustubh, et al.
Publicado: (2025) -
AMD Versal Implementations of FAM and SSCA Estimators
por: Li, Carol Jingyi, et al.
Publicado: (2025) -
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
por: Ohno, Ayumi, et al.
Publicado: (2025) -
Exploring the Versal AI Engine for 3D Gaussian Splatting
por: Shimamura, Kotaro, et al.
Publicado: (2025) -
CAT: Customized Transformer Accelerator Framework on Versal ACAP
por: Zhang, Wenbo, et al.
Publicado: (2024)