MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
Fuente:
arXiv
Guardado en:
| Autores principales: | Jeong, Minki, Jung, Wanyeong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
por: Aimone, James B
Publicado: (2025)
por: Aimone, James B
Publicado: (2025)
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
por: Serpen, Gursel
Publicado: (2025)
por: Serpen, Gursel
Publicado: (2025)
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026)
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
por: Shivdikar, Kaustubh, et al.
Publicado: (2024)
por: Shivdikar, Kaustubh, et al.
Publicado: (2024)
STEMS: Spatial-Temporal Mapping For Spiking Neural Networks
por: Eissa, Sherif, et al.
Publicado: (2025)
por: Eissa, Sherif, et al.
Publicado: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
por: Shen, Aofeng, et al.
Publicado: (2025)
por: Shen, Aofeng, et al.
Publicado: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
por: Mutlu, Onur, et al.
Publicado: (2024)
por: Mutlu, Onur, et al.
Publicado: (2024)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis
por: Yuksel, Ismail Emir, et al.
Publicado: (2024)
por: Yuksel, Ismail Emir, et al.
Publicado: (2024)
Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis
por: Yuksel, Ismail Emir, et al.
Publicado: (2024)
por: Yuksel, Ismail Emir, et al.
Publicado: (2024)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
In-DRAM True Random Number Generation Using Simultaneous Multiple-Row Activation: An Experimental Study of Real DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2025)
por: Yuksel, Ismail Emir, et al.
Publicado: (2025)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
Assessing the Performance of Analog Training for Transfer Learning
por: Fagbohungbe, Omobayode, et al.
Publicado: (2025)
por: Fagbohungbe, Omobayode, et al.
Publicado: (2025)
OpenRASE: Service Function Chain Emulation
por: Krishnamohan, Theviyanthan, et al.
Publicado: (2025)
por: Krishnamohan, Theviyanthan, et al.
Publicado: (2025)
Towards a Decentralised Application-Centric Orchestration Framework in the Cloud-Edge Continuum
por: Ullah, Amjad, et al.
Publicado: (2025)
por: Ullah, Amjad, et al.
Publicado: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
Leveraging SIMD for Accelerating Large-number Arithmetic
por: Das, Subhrajit, et al.
Publicado: (2026)
por: Das, Subhrajit, et al.
Publicado: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025)
por: Asquini, Lorenzo, et al.
Publicado: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
por: Udayashankar, Sreeharsha, et al.
Publicado: (2025)
por: Udayashankar, Sreeharsha, et al.
Publicado: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
por: Zhang, Qijun, et al.
Publicado: (2026)
por: Zhang, Qijun, et al.
Publicado: (2026)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
por: Elwasif, Wael, et al.
Publicado: (2022)
por: Elwasif, Wael, et al.
Publicado: (2022)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
por: Sharma, Harsh, et al.
Publicado: (2023)
por: Sharma, Harsh, et al.
Publicado: (2023)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
por: Prakriya, Neha, et al.
Publicado: (2023)
por: Prakriya, Neha, et al.
Publicado: (2023)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
por: Liu, Xingyu, et al.
Publicado: (2025)
por: Liu, Xingyu, et al.
Publicado: (2025)
Taking Cryptography Out of the Data Path via Near-Memory Processing in DRAM
por: Barcarolo, Nicola, et al.
Publicado: (2026)
por: Barcarolo, Nicola, et al.
Publicado: (2026)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
por: Shi, Man, et al.
Publicado: (2024)
por: Shi, Man, et al.
Publicado: (2024)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024)
por: Cao, Yingqi, et al.
Publicado: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025)
por: Feng, Weigang, et al.
Publicado: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
por: Zou, An, et al.
Publicado: (2025)
por: Zou, An, et al.
Publicado: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
Ejemplares similares
-
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
por: Aimone, James B
Publicado: (2025) -
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
por: Serpen, Gursel
Publicado: (2025) -
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026) -
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
por: Shivdikar, Kaustubh, et al.
Publicado: (2024) -
STEMS: Spatial-Temporal Mapping For Spiking Neural Networks
por: Eissa, Sherif, et al.
Publicado: (2025)