AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Feng, Zhang, Tunhou, Zhang, Junyao, Ku, Jonathan Hao-Cheng, Wang, Yitu, Yang, Xiaoxuan, Hai, Li, Chen, Yiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
von: Wang, Yitu, et al.
Veröffentlicht: (2023)
von: Wang, Yitu, et al.
Veröffentlicht: (2023)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
OffRAC: Offloading Through Remote Accelerator Calls
von: Yang, Ziyi, et al.
Veröffentlicht: (2025)
von: Yang, Ziyi, et al.
Veröffentlicht: (2025)
ModSRAM: Algorithm-Hardware Co-Design for Large Number Modular Multiplication in SRAM
von: Ku, Jonathan, et al.
Veröffentlicht: (2024)
von: Ku, Jonathan, et al.
Veröffentlicht: (2024)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
von: Chen, Peilin, et al.
Veröffentlicht: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
von: Gao, Hanyuan, et al.
Veröffentlicht: (2026)
von: Gao, Hanyuan, et al.
Veröffentlicht: (2026)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
DL-PIM: Improving Data Locality in Processing-in-Memory Systems
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
von: Fu, Fangfa, et al.
Veröffentlicht: (2024)
von: Fu, Fangfa, et al.
Veröffentlicht: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
von: Guo, Cong, et al.
Veröffentlicht: (2025)
von: Guo, Cong, et al.
Veröffentlicht: (2025)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
CRYPTONITE: Scalable Accelerator Design for Cryptographic Primitives and Algorithms
von: Maheswaran, Karthikeya Sharma, et al.
Veröffentlicht: (2025)
von: Maheswaran, Karthikeya Sharma, et al.
Veröffentlicht: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
von: Li, Huize, et al.
Veröffentlicht: (2026)
von: Li, Huize, et al.
Veröffentlicht: (2026)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
Accelerating Time Series Analysis via Processing using Non-Volatile Memories
von: Fernandez, Ivan, et al.
Veröffentlicht: (2022)
von: Fernandez, Ivan, et al.
Veröffentlicht: (2022)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
von: Li, Chongxiao, et al.
Veröffentlicht: (2024)
von: Li, Chongxiao, et al.
Veröffentlicht: (2024)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
Oobleck: Low-Compromise Design for Fault Tolerant Accelerators
von: Wilks, Guy, et al.
Veröffentlicht: (2025)
von: Wilks, Guy, et al.
Veröffentlicht: (2025)
AutoPower: Automated Few-Shot Architecture-Level Power Modeling by Power Group Decoupling
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
von: Zhang, Qijun, et al.
Veröffentlicht: (2025)
ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
von: Chen, Hanqiu, et al.
Veröffentlicht: (2024)
von: Chen, Hanqiu, et al.
Veröffentlicht: (2024)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
AutoINV: Automated Invariant Generation Framework for Formal Verification on High-Level Synthesis Designs
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2026)
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
von: Wang, Chuanzhen, et al.
Veröffentlicht: (2026)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025) -
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
von: Wang, Yitu, et al.
Veröffentlicht: (2023) -
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025) -
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
von: Cheng, Feng, et al.
Veröffentlicht: (2025) -
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
von: Chen, Peilin, et al.
Veröffentlicht: (2025)