Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Wenyong, Liu, Zhengwu, Ren, Yuan, Wong, Ngai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
por: Feng, Yuannuo, et al.
Publicado: (2025)
por: Feng, Yuannuo, et al.
Publicado: (2025)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
por: Feng, Yuannuo, et al.
Publicado: (2025)
por: Feng, Yuannuo, et al.
Publicado: (2025)
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
por: Zhou, Wenyong, et al.
Publicado: (2025)
por: Zhou, Wenyong, et al.
Publicado: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
por: Hooper, Coleman, et al.
Publicado: (2025)
por: Hooper, Coleman, et al.
Publicado: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
por: Kim, Jiyoon, et al.
Publicado: (2025)
por: Kim, Jiyoon, et al.
Publicado: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
por: Chen, Yuzong, et al.
Publicado: (2024)
por: Chen, Yuzong, et al.
Publicado: (2024)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
por: Wu, Taiqiang, et al.
Publicado: (2025)
por: Wu, Taiqiang, et al.
Publicado: (2025)
Exploration of Activation Fault Reliability in Quantized Systolic Array-Based DNN Accelerators
por: Taheri, Mahdi, et al.
Publicado: (2024)
por: Taheri, Mahdi, et al.
Publicado: (2024)
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
por: Choi, Yuseon, et al.
Publicado: (2025)
por: Choi, Yuseon, et al.
Publicado: (2025)
A 65nm 8b-Activation 8b-Weight SRAM-Based Charge-Domain Computing-in-Memory Macro Using A Fully-Parallel Analog Adder Network and A Single-ADC Interface
por: Yin, Guodong, et al.
Publicado: (2022)
por: Yin, Guodong, et al.
Publicado: (2022)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
por: Wolters, Christopher, et al.
Publicado: (2024)
por: Wolters, Christopher, et al.
Publicado: (2024)
Accelerating PoT Quantization on Edge Devices
por: Saha, Rappy, et al.
Publicado: (2024)
por: Saha, Rappy, et al.
Publicado: (2024)
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
por: Chen, Yuzong, et al.
Publicado: (2024)
por: Chen, Yuzong, et al.
Publicado: (2024)
U-SWIM: Universal Selective Write-Verify for Computing-in-Memory Neural Accelerators
por: Yan, Zheyu, et al.
Publicado: (2023)
por: Yan, Zheyu, et al.
Publicado: (2023)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
por: Duan, Bowen, et al.
Publicado: (2026)
por: Duan, Bowen, et al.
Publicado: (2026)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
por: Xie, Xilong, et al.
Publicado: (2025)
por: Xie, Xilong, et al.
Publicado: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
por: Wang, Chenyu, et al.
Publicado: (2023)
por: Wang, Chenyu, et al.
Publicado: (2023)
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
por: Qin, Yifan, et al.
Publicado: (2026)
por: Qin, Yifan, et al.
Publicado: (2026)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
por: Matsushima, Kosuke, et al.
Publicado: (2026)
por: Matsushima, Kosuke, et al.
Publicado: (2026)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
por: Bhattacharya, Swastik, et al.
Publicado: (2025)
por: Bhattacharya, Swastik, et al.
Publicado: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
por: İslamoğlu, Gamze, et al.
Publicado: (2023)
por: İslamoğlu, Gamze, et al.
Publicado: (2023)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
por: AbouElhamayed, Ahmed F., et al.
Publicado: (2023)
por: AbouElhamayed, Ahmed F., et al.
Publicado: (2023)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
por: Dong, Shuai, et al.
Publicado: (2026)
por: Dong, Shuai, et al.
Publicado: (2026)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
por: Doumet, Mario, et al.
Publicado: (2024)
por: Doumet, Mario, et al.
Publicado: (2024)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
por: Klhufek, Jan, et al.
Publicado: (2024)
por: Klhufek, Jan, et al.
Publicado: (2024)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
por: de Lima, João Paulo Cardoso, et al.
Publicado: (2025)
por: de Lima, João Paulo Cardoso, et al.
Publicado: (2025)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
por: Kim, Daeun, et al.
Publicado: (2025)
por: Kim, Daeun, et al.
Publicado: (2025)
Bit-Flip Fault Attack: Crushing Graph Neural Networks via Gradual Bit Search
por: Abharian, Sanaz Kazemi, et al.
Publicado: (2025)
por: Abharian, Sanaz Kazemi, et al.
Publicado: (2025)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
por: Fang, Jiaxun, et al.
Publicado: (2025)
por: Fang, Jiaxun, et al.
Publicado: (2025)
Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
por: Zou, Qingyun, et al.
Publicado: (2026)
por: Zou, Qingyun, et al.
Publicado: (2026)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
por: Yang, Xiaoxuan, et al.
Publicado: (2025)
por: Yang, Xiaoxuan, et al.
Publicado: (2025)
A2Q+: Improving Accumulator-Aware Weight Quantization
por: Colbert, Ian, et al.
Publicado: (2024)
por: Colbert, Ian, et al.
Publicado: (2024)
Kernel Approximation using Analog In-Memory Computing
por: Büchel, Julian, et al.
Publicado: (2024)
por: Büchel, Julian, et al.
Publicado: (2024)
ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
por: Liu, Hongxiang, et al.
Publicado: (2025)
por: Liu, Hongxiang, et al.
Publicado: (2025)
Accelerating Computer Architecture Simulation through Machine Learning
por: Ali, Wajid, et al.
Publicado: (2024)
por: Ali, Wajid, et al.
Publicado: (2024)
FedBit: Accelerating Privacy-Preserving Federated Learning via Bit-Interleaved Packing and Cross-Layer Co-Design
por: Meng, Xiangchen, et al.
Publicado: (2025)
por: Meng, Xiangchen, et al.
Publicado: (2025)
RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer Acceleration
por: Zhao, Lei, et al.
Publicado: (2023)
por: Zhao, Lei, et al.
Publicado: (2023)
Sustainable Transformer Neural Network Acceleration with Stochastic Photonic Computing
por: Afifi, S., et al.
Publicado: (2026)
por: Afifi, S., et al.
Publicado: (2026)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
por: Kim, Wonung, et al.
Publicado: (2025)
por: Kim, Wonung, et al.
Publicado: (2025)
Ejemplares similares
-
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
por: Feng, Yuannuo, et al.
Publicado: (2025) -
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
por: Feng, Yuannuo, et al.
Publicado: (2025) -
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
por: Zhou, Wenyong, et al.
Publicado: (2025) -
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
por: Hooper, Coleman, et al.
Publicado: (2025) -
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
por: Kim, Jiyoon, et al.
Publicado: (2025)