CAMformer: Associative Memory is All You Need
Fuente:
arXiv
Saved in:
| Main Authors: | Molom-Ochir, Tergel, Morris, Benjamin F., Horton, Mark, Wei, Chiyue, Guo, Cong, Taylor, Brady, Liu, Peter, Wang, Shan X., Fan, Deliang, Li, Hai Helen, Chen, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
by: Molom-Ochir, Tergel, et al.
Published: (2024)
by: Molom-Ochir, Tergel, et al.
Published: (2024)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
by: Yang, Xiaoxuan, et al.
Published: (2025)
by: Yang, Xiaoxuan, et al.
Published: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025)
by: Shan, Haoxuan, et al.
Published: (2025)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
by: Fu, Yuzhe, et al.
Published: (2025)
by: Fu, Yuzhe, et al.
Published: (2025)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
by: Wolters, Christopher, et al.
Published: (2024)
by: Wolters, Christopher, et al.
Published: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
by: Gu, Yufeng, et al.
Published: (2025)
by: Gu, Yufeng, et al.
Published: (2025)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
by: Horton, Mark, et al.
Published: (2025)
by: Horton, Mark, et al.
Published: (2025)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
Towards Error Correction for Computing in Racetrack Memory
by: Brazzle, Preston, et al.
Published: (2024)
by: Brazzle, Preston, et al.
Published: (2024)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
by: Wang, Yitu, et al.
Published: (2023)
by: Wang, Yitu, et al.
Published: (2023)
CORDIC Is All You Need
by: Kokane, Omkar, et al.
Published: (2025)
by: Kokane, Omkar, et al.
Published: (2025)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
CAMASim: A Comprehensive Simulation Framework for Content-Addressable Memory based Accelerators
by: Li, Mengyuan, et al.
Published: (2024)
by: Li, Mengyuan, et al.
Published: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
Octopus: Enhancing CXL Memory Pods via Sparse Topology
by: Zhong, Yuhong, et al.
Published: (2025)
by: Zhong, Yuhong, et al.
Published: (2025)
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
by: Hashem, Maeesha Binte, et al.
Published: (2024)
by: Hashem, Maeesha Binte, et al.
Published: (2024)
Error Detection and Correction Codes for Safe In-Memory Computations
by: Parrini, Luca, et al.
Published: (2024)
by: Parrini, Luca, et al.
Published: (2024)
Count2Multiply: Reliable In-Memory High-Radix Counting
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
All-in-Memory Stochastic Computing using ReRAM
by: de Lima, João Paulo C., et al.
Published: (2025)
by: de Lima, João Paulo C., et al.
Published: (2025)
The Case for Replication-Aware Memory-Error Protection in Disaggregated Memory
by: Volos, Haris
Published: (2023)
by: Volos, Haris
Published: (2023)
Exploring and Exploiting Runtime Reconfigurable Floating Point Precision in Scientific Computing: a Case Study for Solving PDEs
by: Hao, Cong "Callie"
Published: (2024)
by: Hao, Cong "Callie"
Published: (2024)
In-Memory ADC-Based Nonlinear Activation Quantization for Efficient In-Memory Computing
by: Dong, Shuai, et al.
Published: (2026)
by: Dong, Shuai, et al.
Published: (2026)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025)
by: Basalama, Suhail, et al.
Published: (2025)
The Future of Memory: Limits and Opportunities
by: Dayo, Samuel, et al.
Published: (2025)
by: Dayo, Samuel, et al.
Published: (2025)
MICSim: A Modular Simulator for Mixed-signal Compute-in-Memory based AI Accelerator
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms
by: Khan, Asif Ali, et al.
Published: (2022)
by: Khan, Asif Ali, et al.
Published: (2022)
PUMA: Efficient and Low-Cost Memory Allocation and Alignment Support for Processing-Using-Memory Architectures
by: Oliveira, Geraldo F., et al.
Published: (2024)
by: Oliveira, Geraldo F., et al.
Published: (2024)
Similar Items
-
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
by: Molom-Ochir, Tergel, et al.
Published: (2024) -
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
by: Yang, Xiaoxuan, et al.
Published: (2025) -
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025) -
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
by: Cheng, Feng, et al.
Published: (2025) -
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
by: Wei, Chiyue, et al.
Published: (2025)