CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Ye, Fang, Chao, Song, Xiaoyong, Wu, Qi, Jiang, Anying, Bai, Yichuan, Du, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
von: Jeon, Sangmin, et al.
Veröffentlicht: (2025)
von: Jeon, Sangmin, et al.
Veröffentlicht: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
Annotated PIM Bibliography
von: Kogge, Peter M.
Veröffentlicht: (2026)
von: Kogge, Peter M.
Veröffentlicht: (2026)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
von: Malekar, Jinendra, et al.
Veröffentlicht: (2025)
von: Malekar, Jinendra, et al.
Veröffentlicht: (2025)
AME-PIM: Can Memory be Your Next Tensor Accelerator?
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
von: Jang, Yongjoo, et al.
Veröffentlicht: (2025)
von: Jang, Yongjoo, et al.
Veröffentlicht: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
von: Kabir, MD Arafat, et al.
Veröffentlicht: (2024)
von: Kabir, MD Arafat, et al.
Veröffentlicht: (2024)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
DL-PIM: Improving Data Locality in Processing-in-Memory Systems
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
GEN-Graph: Heterogeneous PIM Accelerator for General Computational Patterns in Graph-based Dynamic Programming
von: Chen, Yanru, et al.
Veröffentlicht: (2026)
von: Chen, Yanru, et al.
Veröffentlicht: (2026)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM Architecture
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
von: Leitersdorf, Orian, et al.
Veröffentlicht: (2023)
von: Leitersdorf, Orian, et al.
Veröffentlicht: (2023)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
Designing Efficient LLM Accelerators for Edge Devices
von: Haris, Jude, et al.
Veröffentlicht: (2024)
von: Haris, Jude, et al.
Veröffentlicht: (2024)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025) -
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
von: Heo, Guseul, et al.
Veröffentlicht: (2024) -
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
von: Cha, SangHoon, et al.
Veröffentlicht: (2026) -
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
von: Jeon, Sangmin, et al.
Veröffentlicht: (2025) -
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)