Computing-In-Memory Aware Model Adaption For Edge Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Ming-Han, Chang, Tian-Sheuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915969217069056
author Lin, Ming-Han
Chang, Tian-Sheuan
author_facet Lin, Ming-Han
Chang, Tian-Sheuan
contents Computing-in-Memory (CIM) macros have gained popularity for deep learning acceleration due to their highly parallel computation and low power consumption. However, limited macro size and ADC precision introduce throughput and accuracy bottlenecks. This paper proposes a two-stage CIM-aware model adaptation process. The first stage compresses the model and reallocates resources based on layer importance and macro size constraints, reducing model weight loading latency while improving resource utilization and maintaining accuracy. The second stage performs quantization-aware training, incorporating partial sum quantization and ADC precision to mitigate quantization errors in inference. The proposed approach enhances CIM array utilization to 90\%, enables concurrent activation of up to 256 word lines, and achieves up to 93\% compression, all while preserving accuracy comparable to previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Computing-In-Memory Aware Model Adaption For Edge Devices
Lin, Ming-Han
Chang, Tian-Sheuan
Hardware Architecture
Computing-in-Memory (CIM) macros have gained popularity for deep learning acceleration due to their highly parallel computation and low power consumption. However, limited macro size and ADC precision introduce throughput and accuracy bottlenecks. This paper proposes a two-stage CIM-aware model adaptation process. The first stage compresses the model and reallocates resources based on layer importance and macro size constraints, reducing model weight loading latency while improving resource utilization and maintaining accuracy. The second stage performs quantization-aware training, incorporating partial sum quantization and ADC precision to mitigate quantization errors in inference. The proposed approach enhances CIM array utilization to 90\%, enables concurrent activation of up to 256 word lines, and achieves up to 93\% compression, all while preserving accuracy comparable to previous methods.
title Computing-In-Memory Aware Model Adaption For Edge Devices
topic Hardware Architecture
url https://arxiv.org/abs/2510.14379