Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
Fuente:
arXiv
Saved in:
| Main Authors: | Chowdhury, Md Rownak Hossain, Rahman, Mostafizur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating PageRank Algorithmic Tasks with a new Programmable Hardware Architecture
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2024)
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
by: Klhufek, Jan, et al.
Published: (2024)
by: Klhufek, Jan, et al.
Published: (2024)
Energy-Aware Deep Learning on Resource-Constrained Hardware
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
by: Wang, Irene, et al.
Published: (2025)
by: Wang, Irene, et al.
Published: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
by: Hsiung, Ching-Lin, et al.
Published: (2025)
by: Hsiung, Ching-Lin, et al.
Published: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025)
by: Mueller, Lion, et al.
Published: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
by: Andronic, Marta, et al.
Published: (2025)
by: Andronic, Marta, et al.
Published: (2025)
Hardware-Aware Fine-Tuning of Spiking Q-Networks on the SpiNNaker2 Neuromorphic Platform
by: Arfa, Sirine, et al.
Published: (2025)
by: Arfa, Sirine, et al.
Published: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
by: Killian, Earl
Published: (2026)
by: Killian, Earl
Published: (2026)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
by: Li, Enlai, et al.
Published: (2026)
by: Li, Enlai, et al.
Published: (2026)
A Joint Learning Approach to Hardware Caching and Prefetching
by: Yuan, Samuel, et al.
Published: (2025)
by: Yuan, Samuel, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
by: Biggs, Benjamin, et al.
Published: (2023)
by: Biggs, Benjamin, et al.
Published: (2023)
Data-Rate-Aware High-Speed CNN Inference on FPGAs
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series Classification
by: Ikeda, Sosei, et al.
Published: (2025)
by: Ikeda, Sosei, et al.
Published: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Hardware implementation of timely reliable Bayesian decision-making using memristors
by: Song, Lekai, et al.
Published: (2024)
by: Song, Lekai, et al.
Published: (2024)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
by: Zhang, Zehuan, et al.
Published: (2026)
by: Zhang, Zehuan, et al.
Published: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
by: Choi, Dawon, et al.
Published: (2026)
by: Choi, Dawon, et al.
Published: (2026)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
VeriBug: An Attention-based Framework for Bug-Localization in Hardware Designs
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
Mugi: Value Level Parallelism For Efficient LLMs
by: Price, Daniel, et al.
Published: (2026)
by: Price, Daniel, et al.
Published: (2026)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
by: Fyon, Arthur, et al.
Published: (2026)
by: Fyon, Arthur, et al.
Published: (2026)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
by: Xiang, Maoyang, et al.
Published: (2025)
by: Xiang, Maoyang, et al.
Published: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
by: Zhang, Zixi, et al.
Published: (2023)
by: Zhang, Zixi, et al.
Published: (2023)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
by: Tran, Van Duy, et al.
Published: (2024)
by: Tran, Van Duy, et al.
Published: (2024)
Similar Items
-
Accelerating PageRank Algorithmic Tasks with a new Programmable Hardware Architecture
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2024) -
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025) -
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
by: Chowdhury, Md. Rownak Hossain, et al.
Published: (2024) -
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026) -
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
by: Klhufek, Jan, et al.
Published: (2024)