The Role of Advanced Computer Architectures in Accelerating Artificial Intelligence Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Amin, Shahid, Shah, Syed Pervez Hussnain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025)
by: Shukla, Arnav, et al.
Published: (2025)
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025)
by: Prakash, Shvetank, et al.
Published: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
by: Yadav, Divakar Kumar, et al.
Published: (2026)
by: Yadav, Divakar Kumar, et al.
Published: (2026)
The Phantom of PCIe: Constraining Generative Artificial Intelligences for Practical Peripherals Trace Synthesizing
by: Huang, Zhibai, et al.
Published: (2024)
by: Huang, Zhibai, et al.
Published: (2024)
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
by: Helal, Shereef, et al.
Published: (2025)
by: Helal, Shereef, et al.
Published: (2025)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025)
by: Ma, Shaobo, et al.
Published: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
ElasticAI: Creating and Deploying Energy-Efficient Deep Learning Accelerator for Pervasive Computing
by: Qian, Chao, et al.
Published: (2024)
by: Qian, Chao, et al.
Published: (2024)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
by: Qin, Yifan, et al.
Published: (2023)
by: Qin, Yifan, et al.
Published: (2023)
InF-ATPG: Intelligent FFR-Driven ATPG with Advanced Circuit Representation Guided Reinforcement Learning
by: Sun, Bin, et al.
Published: (2025)
by: Sun, Bin, et al.
Published: (2025)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities
by: Read, James, et al.
Published: (2025)
by: Read, James, et al.
Published: (2025)
Heterogeneous Acceleration Pipeline for Recommendation System Training
by: Adnan, Muhammad, et al.
Published: (2022)
by: Adnan, Muhammad, et al.
Published: (2022)
Dynamic Sparse Attention: Access Patterns and Architecture
by: Levy, Noam
Published: (2026)
by: Levy, Noam
Published: (2026)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)
by: Lin, Yujun, et al.
Published: (2025)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
by: Demirkiran, Cansu, et al.
Published: (2023)
by: Demirkiran, Cansu, et al.
Published: (2023)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
by: Liu, Mingju, et al.
Published: (2026)
by: Liu, Mingju, et al.
Published: (2026)
Neuromorphic Computing for Low-Power Artificial Intelligence
by: Katti, Keshava, et al.
Published: (2026)
by: Katti, Keshava, et al.
Published: (2026)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
by: Meng, Chang, et al.
Published: (2026)
by: Meng, Chang, et al.
Published: (2026)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
by: Ahmed, Md Rubel, et al.
Published: (2024)
by: Ahmed, Md Rubel, et al.
Published: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
by: Taheri, Mahdi, et al.
Published: (2024)
by: Taheri, Mahdi, et al.
Published: (2024)
SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN Accelerators
by: Taheri, Mahdi, et al.
Published: (2024)
by: Taheri, Mahdi, et al.
Published: (2024)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
by: Li, Guoyu, et al.
Published: (2025)
by: Li, Guoyu, et al.
Published: (2025)
AIRCHITECT v2: Learning the Hardware Accelerator Design Space through Unified Representations
by: Seo, Jamin, et al.
Published: (2025)
by: Seo, Jamin, et al.
Published: (2025)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
by: Sabih, Muhammad, et al.
Published: (2025)
by: Sabih, Muhammad, et al.
Published: (2025)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
by: Baldi, T., et al.
Published: (2026)
by: Baldi, T., et al.
Published: (2026)
GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
by: Abi-Karam, Stefan, et al.
Published: (2023)
by: Abi-Karam, Stefan, et al.
Published: (2023)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
by: Han, Ruichi, et al.
Published: (2026)
by: Han, Ruichi, et al.
Published: (2026)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
From Fuzzy to Exact: The Halo Architecture for Infinite-Depth Reasoning via Rational Arithmetic
by: Ren, Hansheng
Published: (2026)
by: Ren, Hansheng
Published: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
Automated and Holistic Co-design of Neural Networks and ASICs for Enabling In-Pixel Intelligence
by: Kharel, Shubha R., et al.
Published: (2024)
by: Kharel, Shubha R., et al.
Published: (2024)
Similar Items
-
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025) -
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025) -
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
by: Yadav, Divakar Kumar, et al.
Published: (2026) -
The Phantom of PCIe: Constraining Generative Artificial Intelligences for Practical Peripherals Trace Synthesizing
by: Huang, Zhibai, et al.
Published: (2024) -
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
by: Helal, Shereef, et al.
Published: (2025)