A complete discussion on fully reconfigurable, digital, scalable, graph and sparsity-aware near-memory accelerator for graph neural networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raman, Siddhartha Raman Sundara, John, Lizy, Kulkarni, Jaydeep P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026)
Emerging memory technologies at room/cryogenic temperature
von: Raman, Siddhartha Raman Sundara
Veröffentlicht: (2026)
von: Raman, Siddhartha Raman Sundara
Veröffentlicht: (2026)
Addressing memory bandwidth scalability in vector processors for streaming applications
von: Altayo, Jordi, et al.
Veröffentlicht: (2025)
von: Altayo, Jordi, et al.
Veröffentlicht: (2025)
Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training
von: Wang, Yipeng, et al.
Veröffentlicht: (2025)
von: Wang, Yipeng, et al.
Veröffentlicht: (2025)
The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview
von: Khan, Asif Ali, et al.
Veröffentlicht: (2024)
von: Khan, Asif Ali, et al.
Veröffentlicht: (2024)
TaiBai: A fully programmable brain-inspired processor with topology-aware efficiency
von: Li, Qianpeng, et al.
Veröffentlicht: (2025)
von: Li, Qianpeng, et al.
Veröffentlicht: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
A virtually connected probabilistic computer as a solver for higher-order, densely connected, or reconfigurable combinatorial optimisation problems
von: Searle, Amy J., et al.
Veröffentlicht: (2026)
von: Searle, Amy J., et al.
Veröffentlicht: (2026)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
von: Houshmand, Pouya, et al.
Veröffentlicht: (2024)
von: Houshmand, Pouya, et al.
Veröffentlicht: (2024)
Versatile silicon integrated photonic processor: a reconfigurable solution for next-generation AI clusters
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
von: Zhu, Ying, et al.
Veröffentlicht: (2025)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
DAG-aware Synthesis Orchestration
von: Li, Yingjie, et al.
Veröffentlicht: (2023)
von: Li, Yingjie, et al.
Veröffentlicht: (2023)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
NL-DPE: An Analog In-memory Non-Linear Dot Product Engine for Efficient CNN and LLM Inference
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Virtual memory for real-time systems using hPMP
von: Walluszik, Konrad, et al.
Veröffentlicht: (2025)
von: Walluszik, Konrad, et al.
Veröffentlicht: (2025)
CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory Computing
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
von: Colleman, Steven, et al.
Veröffentlicht: (2024)
von: Colleman, Steven, et al.
Veröffentlicht: (2024)
ARCANE: Adaptive RISC-V Cache Architecture for Near-memory Extensions
von: Petrolo, Vincenzo, et al.
Veröffentlicht: (2025)
von: Petrolo, Vincenzo, et al.
Veröffentlicht: (2025)
Code size reduction by advanced near addressing modes
von: Nuernberger, Kajetan, et al.
Veröffentlicht: (2026)
von: Nuernberger, Kajetan, et al.
Veröffentlicht: (2026)
A methodology to automatically optimize dynamic memory managers applying grammatical evolution
von: Risco-Martín, José L., et al.
Veröffentlicht: (2024)
von: Risco-Martín, José L., et al.
Veröffentlicht: (2024)
Analog Bayesian neural networks are insensitive to the shape of the weight distribution
von: Patel, Ravi G., et al.
Veröffentlicht: (2025)
von: Patel, Ravi G., et al.
Veröffentlicht: (2025)
A Case for Kolmogorov-Arnold Networks in Prefetching: Towards Low-Latency, Generalizable ML-Based Prefetchers
von: Kulkarni, Dhruv, et al.
Veröffentlicht: (2025)
von: Kulkarni, Dhruv, et al.
Veröffentlicht: (2025)
Processing-in-memory for genomics workloads
von: Simon, William Andrew, et al.
Veröffentlicht: (2025)
von: Simon, William Andrew, et al.
Veröffentlicht: (2025)
Accelerating GNN Training through Locality-aware Dropout and Merge
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
von: Ahn, Bas, et al.
Veröffentlicht: (2026)
von: Ahn, Bas, et al.
Veröffentlicht: (2026)
EA4RCA:Efficient AIE accelerator design framework for Regular Communication-Avoiding Algorithm
von: Zhang, W. B., et al.
Veröffentlicht: (2024)
von: Zhang, W. B., et al.
Veröffentlicht: (2024)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
von: Cammarata, Danilo, et al.
Veröffentlicht: (2026)
von: Cammarata, Danilo, et al.
Veröffentlicht: (2026)
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
von: Wang, Yanyun, et al.
Veröffentlicht: (2025)
von: Wang, Yanyun, et al.
Veröffentlicht: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
Topkima-Former: Low-energy, Low-Latency Inference for Transformers using top-k In-memory ADC
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026) -
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026) -
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
von: Raman, Siddhartha Raman Sundara, et al.
Veröffentlicht: (2026) -
Emerging memory technologies at room/cryogenic temperature
von: Raman, Siddhartha Raman Sundara
Veröffentlicht: (2026) -
Addressing memory bandwidth scalability in vector processors for streaming applications
von: Altayo, Jordi, et al.
Veröffentlicht: (2025)