AMD Versal Implementations of FAM and SSCA Estimators
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Carol Jingyi, Wu, Ruilin, Leong, Philip H. W. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026)
by: Mhatre, Kaustubh, et al.
Published: (2026)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
by: Lou, Binglei, et al.
Published: (2025)
by: Lou, Binglei, et al.
Published: (2025)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
CAT: Customized Transformer Accelerator Framework on Versal ACAP
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Exploring the Versal AI Engine for 3D Gaussian Splatting
by: Shimamura, Kotaro, et al.
Published: (2025)
by: Shimamura, Kotaro, et al.
Published: (2025)
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
by: Ohno, Ayumi, et al.
Published: (2025)
by: Ohno, Ayumi, et al.
Published: (2025)
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
by: Dai, Tuo, et al.
Published: (2024)
by: Dai, Tuo, et al.
Published: (2024)
Low-Latency FPGA Control System for Real-Time Neural Network Processing in CCD-Based Trapped-Ion Qubit Measurement
by: Lou, Binglei, et al.
Published: (2025)
by: Lou, Binglei, et al.
Published: (2025)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
by: Li, Guoyu, et al.
Published: (2025)
by: Li, Guoyu, et al.
Published: (2025)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
by: Li, Enlai, et al.
Published: (2026)
by: Li, Enlai, et al.
Published: (2026)
Enabling Mixed criticality applications for the Versal AI-Engines
by: Sprave, Vincent, et al.
Published: (2026)
by: Sprave, Vincent, et al.
Published: (2026)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
by: Rösti, André, et al.
Published: (2025)
by: Rösti, André, et al.
Published: (2025)
VersaQ-3D: A Reconfigurable Accelerator Enabling Feed-Forward and Generalizable 3D Reconstruction via Versatile Quantization
by: Zhang, Yipu, et al.
Published: (2026)
by: Zhang, Yipu, et al.
Published: (2026)
Towards High-Performance Network Coding: FPGA Acceleration With Bounded-value Generators
by: Qing, Jiaxin, et al.
Published: (2025)
by: Qing, Jiaxin, et al.
Published: (2025)
Design and Implementation of BNN-Based Object Detection on FPGA
by: Zhao, Xuyu, et al.
Published: (2026)
by: Zhao, Xuyu, et al.
Published: (2026)
fSEAD: a Composable FPGA-based Streaming Ensemble Anomaly Detection Library
by: Lou, Binglei, et al.
Published: (2024)
by: Lou, Binglei, et al.
Published: (2024)
The Tiny Median Filter: A Small Size, Flexible Arbitrary Percentile Finder Scheme Suitable for FPGA Implementation
by: Wu, Jinyuan
Published: (2024)
by: Wu, Jinyuan
Published: (2024)
Online Learning Extreme Learning Machine with Low-Complexity Predictive Plasticity Rule and FPGA Implementation
by: Zang, Zhenya, et al.
Published: (2025)
by: Zang, Zhenya, et al.
Published: (2025)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
by: Lou, Binglei, et al.
Published: (2024)
by: Lou, Binglei, et al.
Published: (2024)
Implementation of Compute Intensive Algorithms on Software Configurable Processor
by: Ganesha, et al.
Published: (2025)
by: Ganesha, et al.
Published: (2025)
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
by: Mecik, Michael, et al.
Published: (2025)
by: Mecik, Michael, et al.
Published: (2025)
Design and Implementation of Washing Machine HUD Using FPGAs
by: Stites, Norman, et al.
Published: (2025)
by: Stites, Norman, et al.
Published: (2025)
A Power-Efficient Hardware Implementation of L-Mul
by: Chen, Ruiqi, et al.
Published: (2024)
by: Chen, Ruiqi, et al.
Published: (2024)
The Future of Memory: Limits and Opportunities
by: Dayo, Samuel, et al.
Published: (2025)
by: Dayo, Samuel, et al.
Published: (2025)
smallNet: Implementation of a convolutional layer in tiny FPGAs
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
Binary Neural Network Implementation for Handwritten Digit Recognition on FPGA
by: Ertörer, Emir Devlet, et al.
Published: (2025)
by: Ertörer, Emir Devlet, et al.
Published: (2025)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024)
by: Sohn, Gina, et al.
Published: (2024)
Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator
by: Ando, Takuto, et al.
Published: (2025)
by: Ando, Takuto, et al.
Published: (2025)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
Efficient Implementation of an Adaptive Transformer Accelerator for Massive MIMO Outdoor Localization
by: Yaman, Ilayda, et al.
Published: (2026)
by: Yaman, Ilayda, et al.
Published: (2026)
Using a Performance Model to Implement a Superscalar CVA6
by: Allart, Côme, et al.
Published: (2024)
by: Allart, Côme, et al.
Published: (2024)
Design, Implementation and Evaluation of the SVNAPOT Extension on a RISC-V Processor
by: Papadopoulos, Nikolaos-Charalampos, et al.
Published: (2024)
by: Papadopoulos, Nikolaos-Charalampos, et al.
Published: (2024)
A Resource-Driven Approach for Implementing CNNs on FPGAs Using Adaptive IPs
by: Magalhães, Philippe, et al.
Published: (2025)
by: Magalhães, Philippe, et al.
Published: (2025)
A RISC-V SOC for Terahertz IoT Devices: Implementation and design challenges
by: Zhong, Xinchao, et al.
Published: (2024)
by: Zhong, Xinchao, et al.
Published: (2024)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
by: Pu, Huanzhi, et al.
Published: (2025)
by: Pu, Huanzhi, et al.
Published: (2025)
Advancing Cloud Computing Capabilities on gem5 by Implementing the RISC-V Hypervisor Extension
by: Fragkoulis, George-Marios, et al.
Published: (2024)
by: Fragkoulis, George-Marios, et al.
Published: (2024)
Similar Items
-
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026) -
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025) -
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
by: Lou, Binglei, et al.
Published: (2025) -
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
by: Lou, Binglei, et al.
Published: (2026) -
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)