Saved in:
| Main Authors: | Li, Guoyu, Zheng, Pengbo, Weng, Jian, Yang, Enshan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.11441 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
CAT: Customized Transformer Accelerator Framework on Versal ACAP
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
by: Dai, Tuo, et al.
Published: (2024)
by: Dai, Tuo, et al.
Published: (2024)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
by: Li, Enlai, et al.
Published: (2026)
by: Li, Enlai, et al.
Published: (2026)
ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
by: Zhang, Qijun, et al.
Published: (2025)
by: Zhang, Qijun, et al.
Published: (2025)
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
by: Curzel, Serena, et al.
Published: (2023)
by: Curzel, Serena, et al.
Published: (2023)
Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and Optimization
by: Blasberg, Alexander, et al.
Published: (2026)
by: Blasberg, Alexander, et al.
Published: (2026)
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
by: Bougioukou, Eleni, et al.
Published: (2025)
by: Bougioukou, Eleni, et al.
Published: (2025)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Architectural Design and Performance Analysis of FPGA based AI Accelerators: A Comprehensive Review
by: Chatterjee, Soumita, et al.
Published: (2026)
by: Chatterjee, Soumita, et al.
Published: (2026)
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
by: Deng, Jinyi, et al.
Published: (2024)
by: Deng, Jinyi, et al.
Published: (2024)
iDSE: Navigating Design Space Exploration in High-Level Synthesis Using LLMs
by: Li, Runkai, et al.
Published: (2025)
by: Li, Runkai, et al.
Published: (2025)
HAPM -- Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
by: Zhang, Jinsong, et al.
Published: (2025)
by: Zhang, Jinsong, et al.
Published: (2025)
MARCA: Mamba Accelerator with ReConfigurable Architecture
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
by: Lee, Jonghun, et al.
Published: (2026)
by: Lee, Jonghun, et al.
Published: (2026)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026)
by: Ma, Songchen, et al.
Published: (2026)
At the Edge of the Heart: ULP FPGA-Based CNN for On-Device Cardiac Feature Extraction in Smart Health Sensors for Astronauts
by: Rahman, Kazi Mohammad Abidur, et al.
Published: (2026)
by: Rahman, Kazi Mohammad Abidur, et al.
Published: (2026)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
by: Li, Guoyu, et al.
Published: (2025)
by: Li, Guoyu, et al.
Published: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
by: Wan, Zishen, et al.
Published: (2024)
by: Wan, Zishen, et al.
Published: (2024)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
by: Zirui, Ma, et al.
Published: (2026)
by: Zirui, Ma, et al.
Published: (2026)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
by: Chen, Yao, et al.
Published: (2022)
by: Chen, Yao, et al.
Published: (2022)
AMD Versal Implementations of FAM and SSCA Estimators
by: Li, Carol Jingyi, et al.
Published: (2025)
by: Li, Carol Jingyi, et al.
Published: (2025)
A 10.60 $μ$W 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
by: Qin, Yifan, et al.
Published: (2024)
by: Qin, Yifan, et al.
Published: (2024)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
by: Abi-Karam, Stefan, et al.
Published: (2025)
by: Abi-Karam, Stefan, et al.
Published: (2025)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
by: Zhong, Ruizhe, et al.
Published: (2023)
by: Zhong, Ruizhe, et al.
Published: (2023)
AnalogSeeker: An Open-source Foundation Language Model for Analog Circuit Design
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
by: Fu, Yuzhe, et al.
Published: (2025)
by: Fu, Yuzhe, et al.
Published: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference
by: Oztas, Ali Emre, et al.
Published: (2026)
by: Oztas, Ali Emre, et al.
Published: (2026)
Assessing Large Language Models in Generating RTL Design Specifications
by: Huang, Hung-Ming, et al.
Published: (2025)
by: Huang, Hung-Ming, et al.
Published: (2025)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
NOVA: NoC-based Vector Unit for Mapping Attention Layers on a CNN Accelerator
by: Upadhyay, Mohit, et al.
Published: (2024)
by: Upadhyay, Mohit, et al.
Published: (2024)
YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
by: Xuan, Zihao, et al.
Published: (2023)
by: Xuan, Zihao, et al.
Published: (2023)
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
by: Zuo, Dongsheng, et al.
Published: (2025)
by: Zuo, Dongsheng, et al.
Published: (2025)
Design Rules for Extreme-Edge Scientific Computing on AI Engines
by: Ma, Zhenghua, et al.
Published: (2026)
by: Ma, Zhenghua, et al.
Published: (2026)
ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
by: Xing, Changwen, et al.
Published: (2025)
by: Xing, Changwen, et al.
Published: (2025)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
Similar Items
-
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025) -
CAT: Customized Transformer Accelerator Framework on Versal ACAP
by: Zhang, Wenbo, et al.
Published: (2024) -
WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on the Versal ACAP Architecture
by: Dai, Tuo, et al.
Published: (2024) -
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025) -
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
by: Li, Enlai, et al.
Published: (2026)