HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
Fuente:
arXiv
Salvato in:
| Autori principali: | Rhee, Myunghyun, Sim, Joonseop, Ahn, Taeyoung, Lee, Seungyong, Yoon, Daegun, Kim, Euiseok, Park, Kyoung, Joo, Youngpyo, Kim, Hoshik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026)
di: Chung, Euijun, et al.
Pubblicazione: (2026)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
di: Wang, Xi, et al.
Pubblicazione: (2025)
di: Wang, Xi, et al.
Pubblicazione: (2025)
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
di: Ruiz, Christian Guzman, et al.
Pubblicazione: (2024)
di: Ruiz, Christian Guzman, et al.
Pubblicazione: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference
di: Oztas, Ali Emre, et al.
Pubblicazione: (2026)
di: Oztas, Ali Emre, et al.
Pubblicazione: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
di: Li, Shiju, et al.
Pubblicazione: (2025)
di: Li, Shiju, et al.
Pubblicazione: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
di: Pan, Xiurui, et al.
Pubblicazione: (2024)
di: Pan, Xiurui, et al.
Pubblicazione: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
di: Kim, Hyeseong, et al.
Pubblicazione: (2026)
CuLifter: Lifting GPU Binaries to Typed IR
di: Zhao, Jisheng, et al.
Pubblicazione: (2026)
di: Zhao, Jisheng, et al.
Pubblicazione: (2026)
VR-Pipe: Streamlining Hardware Graphics Pipeline for Volume Rendering
di: Lee, Junseo, et al.
Pubblicazione: (2025)
di: Lee, Junseo, et al.
Pubblicazione: (2025)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
di: Kim, Dowon, et al.
Pubblicazione: (2025)
di: Kim, Dowon, et al.
Pubblicazione: (2025)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
On the Thermal Vulnerability of 3D-Stacked High-Bandwidth Memory Architectures
di: Elahi, Mehdi, et al.
Pubblicazione: (2025)
di: Elahi, Mehdi, et al.
Pubblicazione: (2025)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
di: Yang, Jinho, et al.
Pubblicazione: (2025)
di: Yang, Jinho, et al.
Pubblicazione: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
di: Noh, Seock-Hwan, et al.
Pubblicazione: (2025)
di: Noh, Seock-Hwan, et al.
Pubblicazione: (2025)
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
di: Yoon, Daegun, et al.
Pubblicazione: (2024)
di: Yoon, Daegun, et al.
Pubblicazione: (2024)
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
di: Kim, Youngsuk, et al.
Pubblicazione: (2024)
di: Kim, Youngsuk, et al.
Pubblicazione: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
di: Kim, Hansung, et al.
Pubblicazione: (2024)
di: Kim, Hansung, et al.
Pubblicazione: (2024)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
di: Elwasif, Wael, et al.
Pubblicazione: (2022)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2024)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
di: Pu, Huanzhi, et al.
Pubblicazione: (2025)
di: Pu, Huanzhi, et al.
Pubblicazione: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
di: Shen, Diyou, et al.
Pubblicazione: (2025)
di: Shen, Diyou, et al.
Pubblicazione: (2025)
GigaAPI for GPU Parallelization
di: Suvarna, M., et al.
Pubblicazione: (2025)
di: Suvarna, M., et al.
Pubblicazione: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
di: Yang, Yiwei, et al.
Pubblicazione: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
Parallelizing a modern GPU simulator
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
Cost-effective Deep Learning Infrastructure with NVIDIA GPU
di: Ghimire, Aatiz, et al.
Pubblicazione: (2025)
di: Ghimire, Aatiz, et al.
Pubblicazione: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026) -
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
di: Wang, Xi, et al.
Pubblicazione: (2025) -
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
di: Ruiz, Christian Guzman, et al.
Pubblicazione: (2024) -
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)