A complete discussion on fully reconfigurable, digital, scalable, graph and sparsity-aware near-memory accelerator for graph neural networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Raman, Siddhartha Raman Sundara, John, Lizy, Kulkarni, Jaydeep P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026)
Emerging memory technologies at room/cryogenic temperature
por: Raman, Siddhartha Raman Sundara
Publicado: (2026)
por: Raman, Siddhartha Raman Sundara
Publicado: (2026)
Addressing memory bandwidth scalability in vector processors for streaming applications
por: Altayo, Jordi, et al.
Publicado: (2025)
por: Altayo, Jordi, et al.
Publicado: (2025)
Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training
por: Wang, Yipeng, et al.
Publicado: (2025)
por: Wang, Yipeng, et al.
Publicado: (2025)
The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview
por: Khan, Asif Ali, et al.
Publicado: (2024)
por: Khan, Asif Ali, et al.
Publicado: (2024)
TaiBai: A fully programmable brain-inspired processor with topology-aware efficiency
por: Li, Qianpeng, et al.
Publicado: (2025)
por: Li, Qianpeng, et al.
Publicado: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
por: Ma, Siyuan, et al.
Publicado: (2025)
por: Ma, Siyuan, et al.
Publicado: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
por: Li, Ruihao, et al.
Publicado: (2026)
por: Li, Ruihao, et al.
Publicado: (2026)
A virtually connected probabilistic computer as a solver for higher-order, densely connected, or reconfigurable combinatorial optimisation problems
por: Searle, Amy J., et al.
Publicado: (2026)
por: Searle, Amy J., et al.
Publicado: (2026)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
por: Juneja, Rohan, et al.
Publicado: (2025)
por: Juneja, Rohan, et al.
Publicado: (2025)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
por: Houshmand, Pouya, et al.
Publicado: (2024)
por: Houshmand, Pouya, et al.
Publicado: (2024)
Versatile silicon integrated photonic processor: a reconfigurable solution for next-generation AI clusters
por: Zhu, Ying, et al.
Publicado: (2025)
por: Zhu, Ying, et al.
Publicado: (2025)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
por: Li, Wanqian, et al.
Publicado: (2024)
por: Li, Wanqian, et al.
Publicado: (2024)
DAG-aware Synthesis Orchestration
por: Li, Yingjie, et al.
Publicado: (2023)
por: Li, Yingjie, et al.
Publicado: (2023)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
NL-DPE: An Analog In-memory Non-Linear Dot Product Engine for Efficient CNN and LLM Inference
por: Zhao, Lei, et al.
Publicado: (2025)
por: Zhao, Lei, et al.
Publicado: (2025)
Virtual memory for real-time systems using hPMP
por: Walluszik, Konrad, et al.
Publicado: (2025)
por: Walluszik, Konrad, et al.
Publicado: (2025)
CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory Computing
por: Dong, Shuai, et al.
Publicado: (2025)
por: Dong, Shuai, et al.
Publicado: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
por: Verhelst, Marian, et al.
Publicado: (2025)
por: Verhelst, Marian, et al.
Publicado: (2025)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
por: Colleman, Steven, et al.
Publicado: (2024)
por: Colleman, Steven, et al.
Publicado: (2024)
ARCANE: Adaptive RISC-V Cache Architecture for Near-memory Extensions
por: Petrolo, Vincenzo, et al.
Publicado: (2025)
por: Petrolo, Vincenzo, et al.
Publicado: (2025)
Code size reduction by advanced near addressing modes
por: Nuernberger, Kajetan, et al.
Publicado: (2026)
por: Nuernberger, Kajetan, et al.
Publicado: (2026)
A methodology to automatically optimize dynamic memory managers applying grammatical evolution
por: Risco-Martín, José L., et al.
Publicado: (2024)
por: Risco-Martín, José L., et al.
Publicado: (2024)
Analog Bayesian neural networks are insensitive to the shape of the weight distribution
por: Patel, Ravi G., et al.
Publicado: (2025)
por: Patel, Ravi G., et al.
Publicado: (2025)
A Case for Kolmogorov-Arnold Networks in Prefetching: Towards Low-Latency, Generalizable ML-Based Prefetchers
por: Kulkarni, Dhruv, et al.
Publicado: (2025)
por: Kulkarni, Dhruv, et al.
Publicado: (2025)
Processing-in-memory for genomics workloads
por: Simon, William Andrew, et al.
Publicado: (2025)
por: Simon, William Andrew, et al.
Publicado: (2025)
Accelerating GNN Training through Locality-aware Dropout and Merge
por: Sun, Gongjian, et al.
Publicado: (2025)
por: Sun, Gongjian, et al.
Publicado: (2025)
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
por: Ahn, Bas, et al.
Publicado: (2026)
por: Ahn, Bas, et al.
Publicado: (2026)
EA4RCA:Efficient AIE accelerator design framework for Regular Communication-Avoiding Algorithm
por: Zhang, W. B., et al.
Publicado: (2024)
por: Zhang, W. B., et al.
Publicado: (2024)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
por: Cammarata, Danilo, et al.
Publicado: (2026)
por: Cammarata, Danilo, et al.
Publicado: (2026)
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
por: Wang, Yanyun, et al.
Publicado: (2025)
por: Wang, Yanyun, et al.
Publicado: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
por: Moitra, Abhishek, et al.
Publicado: (2024)
por: Moitra, Abhishek, et al.
Publicado: (2024)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
por: Zhang, Chenguang, et al.
Publicado: (2024)
por: Zhang, Chenguang, et al.
Publicado: (2024)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
por: Chong, Yue Jiet, et al.
Publicado: (2025)
por: Chong, Yue Jiet, et al.
Publicado: (2025)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
por: Espinosa, Elena, et al.
Publicado: (2025)
por: Espinosa, Elena, et al.
Publicado: (2025)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
por: Zhao, Shixin, et al.
Publicado: (2025)
por: Zhao, Shixin, et al.
Publicado: (2025)
Topkima-Former: Low-energy, Low-Latency Inference for Transformers using top-k In-memory ADC
por: Dong, Shuai, et al.
Publicado: (2024)
por: Dong, Shuai, et al.
Publicado: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
por: Cheng, Feng, et al.
Publicado: (2025)
por: Cheng, Feng, et al.
Publicado: (2025)
Ejemplares similares
-
A comprehensive study on ILP acceleration accounting for sparsity, area, energy, data movement using near-memory architecture
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026) -
ABI: A tightly integrated, unified, sparsity-aware, reconfigurable, compute near-register file/cache GPU architecture with light-weight softmax for deep learning, linear algebra, and Ising compute
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026) -
A comparative study on power delivery aspects of compute-in/near-memory approaches using DRAM
por: Raman, Siddhartha Raman Sundara, et al.
Publicado: (2026) -
Emerging memory technologies at room/cryogenic temperature
por: Raman, Siddhartha Raman Sundara
Publicado: (2026) -
Addressing memory bandwidth scalability in vector processors for streaming applications
por: Altayo, Jordi, et al.
Publicado: (2025)