GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Arghadip, Kundu, Shamik, Raha, Arnab, Ghosh, Soumendu, Mathaikutty, Deepak, Raghunathan, Vijay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
XAMBA: Enabling Efficient State Space Models on Resource-Constrained Neural Processing Units
by: Das, Arghadip, et al.
Published: (2025)
by: Das, Arghadip, et al.
Published: (2025)
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
by: Raha, Arnab, et al.
Published: (2024)
by: Raha, Arnab, et al.
Published: (2024)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
by: Bhattacharya, Swastik, et al.
Published: (2025)
by: Bhattacharya, Swastik, et al.
Published: (2025)
ReGate: Enabling Power Gating in Neural Processing Units
by: Xue, Yuqi, et al.
Published: (2025)
by: Xue, Yuqi, et al.
Published: (2025)
TYTAN: Taylor-series based Non-Linear Activation Engine for Deep Learning Accelerators
by: Pramanik, Soham, et al.
Published: (2025)
by: Pramanik, Soham, et al.
Published: (2025)
Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
by: Rösti, André, et al.
Published: (2025)
by: Rösti, André, et al.
Published: (2025)
GraCo -- A Graph Composer for Integrated Circuits
by: Uhlich, Stefan, et al.
Published: (2024)
by: Uhlich, Stefan, et al.
Published: (2024)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
SiTe CiM: Signed Ternary Computing-in-Memory for Ultra-Low Precision Deep Neural Networks
by: Thakuria, Niharika, et al.
Published: (2024)
by: Thakuria, Niharika, et al.
Published: (2024)
Low-Cost IoT-Enabled Tele-ECG Monitoring for Resource-Constrained Settings: System Design and Prototype
by: Neupane, Seemron, et al.
Published: (2026)
by: Neupane, Seemron, et al.
Published: (2026)
Dynamic Power Control in a Hardware Neural Network with Error-Configurable MAC Units
by: Ghaderi, Maedeh, et al.
Published: (2024)
by: Ghaderi, Maedeh, et al.
Published: (2024)
Towards Efficient Design Verification -- Constrained Random Verification using PyUVM
by: Gadde, Deepak Narayan, et al.
Published: (2024)
by: Gadde, Deepak Narayan, et al.
Published: (2024)
RTGPU: Real-Time Computing with Graphics Processing Units
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
by: Gheibi-Fetrat, Atiyeh, et al.
Published: (2025)
Reuse and Blend: Energy-Efficient Optical Neural Network Enabled by Weight Sharing
by: Xu, Bo, et al.
Published: (2024)
by: Xu, Bo, et al.
Published: (2024)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
by: Durvasula, Sankeerth, et al.
Published: (2024)
by: Durvasula, Sankeerth, et al.
Published: (2024)
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
SAT-based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
by: Tirelli, Cristian, et al.
Published: (2024)
by: Tirelli, Cristian, et al.
Published: (2024)
Accelerating Neural Networks for Large Language Models and Graph Processing with Silicon Photonics
by: Afifi, Salma, et al.
Published: (2024)
by: Afifi, Salma, et al.
Published: (2024)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
by: Gerlinghoff, Daniel, et al.
Published: (2024)
by: Gerlinghoff, Daniel, et al.
Published: (2024)
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
by: Feng, Dahu, et al.
Published: (2025)
by: Feng, Dahu, et al.
Published: (2025)
FsimNNs: An Open-Source Graph Neural Network Platform for SEU Simulation-based Fault Injection
by: Lu, Li, et al.
Published: (2025)
by: Lu, Li, et al.
Published: (2025)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
by: Puigdemont, Pol, et al.
Published: (2024)
by: Puigdemont, Pol, et al.
Published: (2024)
Leveraging Application-Specific Knowledge for Energy-Efficient Deep Learning Accelerators on Resource-Constrained FPGAs
by: Qian, Chao
Published: (2025)
by: Qian, Chao
Published: (2025)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Containerized In-Storage Processing and Computing-Enabled SSD Disaggregation
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
Effective Design Verification -- Constrained Random with Python and Cocotb
by: Gadde, Deepak Narayan, et al.
Published: (2024)
by: Gadde, Deepak Narayan, et al.
Published: (2024)
e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
by: Machetti, Simone, et al.
Published: (2025)
by: Machetti, Simone, et al.
Published: (2025)
A Scalable Resource Management Layer for FPGA SoCs in 6G Radio Units
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
by: Bartzoudis, Nikolaos, et al.
Published: (2025)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
SA-DS: A Dataset for Large Language Model-Driven AI Accelerator Design Generation
by: Vungarala, Deepak, et al.
Published: (2024)
by: Vungarala, Deepak, et al.
Published: (2024)
RPU -- A Reasoning Processing Unit
by: Adiletta, Matthew, et al.
Published: (2026)
by: Adiletta, Matthew, et al.
Published: (2026)
Similar Items
-
XAMBA: Enabling Efficient State Space Models on Resource-Constrained Neural Processing Units
by: Das, Arghadip, et al.
Published: (2025) -
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
by: Raha, Arnab, et al.
Published: (2024) -
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
by: Parvathy, Aradhana Mohan, et al.
Published: (2026) -
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025) -
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)