PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Lian, Zhao, Shixin, Zhou, Yutian, He, Yintao, Wang, Mengdi, Han, Yinhe, Wang, Ying |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
di: Wang, Zeke, et al.
Pubblicazione: (2025)
di: Wang, Zeke, et al.
Pubblicazione: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2022)
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2022)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
A Modern Primer on Processing in Memory
di: Mutlu, Onur, et al.
Pubblicazione: (2020)
di: Mutlu, Onur, et al.
Pubblicazione: (2020)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
di: Ibrahim, Mohamed Assem, et al.
Pubblicazione: (2024)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
di: Mutlu, Onur, et al.
Pubblicazione: (2024)
di: Mutlu, Onur, et al.
Pubblicazione: (2024)
Efficient Architecture for RISC-V Vector Memory Access
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
di: Oliveira, Geraldo F.
Pubblicazione: (2025)
di: Oliveira, Geraldo F.
Pubblicazione: (2025)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
di: Barkhordar, Marzieh, et al.
Pubblicazione: (2026)
di: Barkhordar, Marzieh, et al.
Pubblicazione: (2026)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System
di: He, Yintao, et al.
Pubblicazione: (2025)
di: He, Yintao, et al.
Pubblicazione: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
di: Wang, Jing, et al.
Pubblicazione: (2025)
di: Wang, Jing, et al.
Pubblicazione: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
Memory-Centric Computing: Solving Computing's Memory Problem
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2024)
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
di: Volos, Haris, et al.
Pubblicazione: (2025)
di: Volos, Haris, et al.
Pubblicazione: (2025)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
di: Frouzakis, Manos, et al.
Pubblicazione: (2025)
di: Frouzakis, Manos, et al.
Pubblicazione: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
PIUMA: Programmable Integrated Unified Memory Architecture
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
di: Li, Jiamin, et al.
Pubblicazione: (2025)
di: Li, Jiamin, et al.
Pubblicazione: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
di: Psistakis, Antonis
Pubblicazione: (2025)
di: Psistakis, Antonis
Pubblicazione: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
di: Pan, Lunshuai, et al.
Pubblicazione: (2024)
di: Pan, Lunshuai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026) -
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026) -
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024) -
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)