Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Lian, Zhao, Shixin, Li, Bing, Ren, Haimeng, Xu, Zhaohui, Wang, Mengdi, Li, Xiaowei, Han, Yinhe, Wang, Ying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
COMET: Towards Partical W4A4KV4 LLMs Serving
von: Liu, Lian, et al.
Veröffentlicht: (2024)
von: Liu, Lian, et al.
Veröffentlicht: (2024)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
ChipGPT: How far are we from natural language hardware design
von: Chang, Kaiyan, et al.
Veröffentlicht: (2023)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2023)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
von: Yu, Jinxin, et al.
Veröffentlicht: (2026)
von: Yu, Jinxin, et al.
Veröffentlicht: (2026)
CLASS: A Controller-Centric Layout Synthesizer for Dynamic Quantum Circuits
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
von: Zou, Guoqiang, et al.
Veröffentlicht: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
von: Cai, Siyang, et al.
Veröffentlicht: (2026)
von: Cai, Siyang, et al.
Veröffentlicht: (2026)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
HLSPilot: LLM-based High-Level Synthesis
von: Xiong, Chenwei, et al.
Veröffentlicht: (2024)
von: Xiong, Chenwei, et al.
Veröffentlicht: (2024)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
Affordable HPC: Leveraging Small Clusters for Big Data and Graph Computing
von: Wu, Ruilong, et al.
Veröffentlicht: (2024)
von: Wu, Ruilong, et al.
Veröffentlicht: (2024)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
A Systematic Characterization of LLM Inference on GPUs
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Natural language is not enough: Benchmarking multi-modal generative AI for Verilog generation
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2024)
Multiport Support for Vortex OpenGPU Memory Hierarchy
von: Shin, Injae, et al.
Veröffentlicht: (2025)
von: Shin, Injae, et al.
Veröffentlicht: (2025)
CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators
von: Qu, Songyun, et al.
Veröffentlicht: (2024)
von: Qu, Songyun, et al.
Veröffentlicht: (2024)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
von: Li, Huize, et al.
Veröffentlicht: (2026)
von: Li, Huize, et al.
Veröffentlicht: (2026)
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
von: Chen, Guoci, et al.
Veröffentlicht: (2026)
von: Chen, Guoci, et al.
Veröffentlicht: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024)
von: Luo, Weile, et al.
Veröffentlicht: (2024)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
von: Wang, Jiapin, et al.
Veröffentlicht: (2024)
von: Wang, Jiapin, et al.
Veröffentlicht: (2024)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning
von: Chen, Zhirong, et al.
Veröffentlicht: (2025)
von: Chen, Zhirong, et al.
Veröffentlicht: (2025)
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
von: Han, Dengke, et al.
Veröffentlicht: (2025)
von: Han, Dengke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026) -
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025) -
COMET: Towards Partical W4A4KV4 LLMs Serving
von: Liu, Lian, et al.
Veröffentlicht: (2024) -
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026) -
ChipGPT: How far are we from natural language hardware design
von: Chang, Kaiyan, et al.
Veröffentlicht: (2023)