HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
Fuente:
arXiv
Saved in:
| Main Authors: | Rhee, Myunghyun, Sim, Joonseop, Ahn, Taeyoung, Lee, Seungyong, Yoon, Daegun, Kim, Euiseok, Park, Kyoung, Joo, Youngpyo, Kim, Hoshik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025)
by: Rhee, Myunghyun, et al.
Published: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
by: Ruiz, Christian Guzman, et al.
Published: (2024)
by: Ruiz, Christian Guzman, et al.
Published: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
by: Yoon, Jieon, et al.
Published: (2026)
by: Yoon, Jieon, et al.
Published: (2026)
DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference
by: Oztas, Ali Emre, et al.
Published: (2026)
by: Oztas, Ali Emre, et al.
Published: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024)
by: Pan, Xiurui, et al.
Published: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
CuLifter: Lifting GPU Binaries to Typed IR
by: Zhao, Jisheng, et al.
Published: (2026)
by: Zhao, Jisheng, et al.
Published: (2026)
VR-Pipe: Streamlining Hardware Graphics Pipeline for Volume Rendering
by: Lee, Junseo, et al.
Published: (2025)
by: Lee, Junseo, et al.
Published: (2025)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
by: Moon, Seungjae, et al.
Published: (2024)
by: Moon, Seungjae, et al.
Published: (2024)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
by: Sharma, Debendra Das, et al.
Published: (2025)
by: Sharma, Debendra Das, et al.
Published: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
On the Thermal Vulnerability of 3D-Stacked High-Bandwidth Memory Architectures
by: Elahi, Mehdi, et al.
Published: (2025)
by: Elahi, Mehdi, et al.
Published: (2025)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
by: Yang, Jinho, et al.
Published: (2025)
by: Yang, Jinho, et al.
Published: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
by: Yoon, Daegun, et al.
Published: (2024)
by: Yoon, Daegun, et al.
Published: (2024)
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
by: Kim, Youngsuk, et al.
Published: (2024)
by: Kim, Youngsuk, et al.
Published: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
by: Elwasif, Wael, et al.
Published: (2022)
by: Elwasif, Wael, et al.
Published: (2022)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
by: Pu, Huanzhi, et al.
Published: (2025)
by: Pu, Huanzhi, et al.
Published: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
by: Shen, Diyou, et al.
Published: (2025)
by: Shen, Diyou, et al.
Published: (2025)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
by: Yang, Yiwei, et al.
Published: (2025)
by: Yang, Yiwei, et al.
Published: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Cost-effective Deep Learning Infrastructure with NVIDIA GPU
by: Ghimire, Aatiz, et al.
Published: (2025)
by: Ghimire, Aatiz, et al.
Published: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
Similar Items
-
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
by: Rhee, Myunghyun, et al.
Published: (2025) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026) -
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025) -
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
by: Ruiz, Christian Guzman, et al.
Published: (2024) -
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
by: Lee, Hangyeol, et al.
Published: (2026)