YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
Fuente:
arXiv
Saved in:
| Main Authors: | Xuan, Zihao, Yang, Yuxuan, Xuan, Wei, Su, Zijia, Chen, Song, Kang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
Implementation of a 8-bit Wallace Tree Multiplier
by: Biswas, Ayan, et al.
Published: (2025)
by: Biswas, Ayan, et al.
Published: (2025)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
A 33.6-136.2 TOPS/W Nonlinear Analog Computing-In-Memory Macro for Multi-bit LSTM Accelerator in 65 nm CMOS
by: Yang, Junyi, et al.
Published: (2025)
by: Yang, Junyi, et al.
Published: (2025)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
by: Ke, Chih-Hua
Published: (2026)
by: Ke, Chih-Hua
Published: (2026)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
WebRISC-V: A 64-bit RISC-V Pipeline Simulator for Computer Architecture Classes
by: Giorgi, Roberto, et al.
Published: (2025)
by: Giorgi, Roberto, et al.
Published: (2025)
Modulo-$(2^{2n}+1)$ Arithmetic via Two Parallel n-bit Residue Channels
by: Jaberipur, Ghassem, et al.
Published: (2024)
by: Jaberipur, Ghassem, et al.
Published: (2024)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
by: Chakraborty, Subhradip, et al.
Published: (2026)
by: Chakraborty, Subhradip, et al.
Published: (2026)
In-Memory Computing Architecture for Efficient Hardware Security
by: Ajmi, Hala, et al.
Published: (2024)
by: Ajmi, Hala, et al.
Published: (2024)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
Architectural Limits of Cloud TPUs in Finite-Field Cryptography
by: Dang, Hung, et al.
Published: (2026)
by: Dang, Hung, et al.
Published: (2026)
Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations
by: Heidary, Masoud, et al.
Published: (2026)
by: Heidary, Masoud, et al.
Published: (2026)
A Hybrid-Domain Floating-Point Compute-in-Memory Architecture for Efficient Acceleration of High-Precision Deep Neural Networks
by: Yi, Zhiqiang, et al.
Published: (2025)
by: Yi, Zhiqiang, et al.
Published: (2025)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
by: Song, Chang Eun, et al.
Published: (2025)
by: Song, Chang Eun, et al.
Published: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
by: Gowda, Bindu G, et al.
Published: (2025)
by: Gowda, Bindu G, et al.
Published: (2025)
Design of a Reformed Array Logic Binary Multiplier for High-Speed Computations
by: Mohammad, Sakib, et al.
Published: (2024)
by: Mohammad, Sakib, et al.
Published: (2024)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
by: Antunes, Pedro, et al.
Published: (2026)
by: Antunes, Pedro, et al.
Published: (2026)
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
by: Vafaei, Jafar, et al.
Published: (2024)
by: Vafaei, Jafar, et al.
Published: (2024)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
by: Wu, Jiajun, et al.
Published: (2024)
by: Wu, Jiajun, et al.
Published: (2024)
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
by: Xuan, Wei, et al.
Published: (2025)
by: Xuan, Wei, et al.
Published: (2025)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
by: Sestito, Cristian, et al.
Published: (2025)
by: Sestito, Cristian, et al.
Published: (2025)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
by: Spagnolo, Tommaso, et al.
Published: (2026)
by: Spagnolo, Tommaso, et al.
Published: (2026)
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
by: Zhang, Wenlun, et al.
Published: (2024)
by: Zhang, Wenlun, et al.
Published: (2024)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
Count2Multiply: Reliable In-Memory High-Radix Counting
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
A Novel 8T SRAM-Based In-Memory Computing Architecture for MAC-Derived Logical Functions
by: M, Amogh K, et al.
Published: (2025)
by: M, Amogh K, et al.
Published: (2025)
Similar Items
-
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026) -
Implementation of a 8-bit Wallace Tree Multiplier
by: Biswas, Ayan, et al.
Published: (2025) -
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024) -
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025) -
A 33.6-136.2 TOPS/W Nonlinear Analog Computing-In-Memory Macro for Multi-bit LSTM Accelerator in 65 nm CMOS
by: Yang, Junyi, et al.
Published: (2025)