A Systematic Characterization of LLM Inference on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haonan, Xiao, Xuxin, Yan, Mingyu, Zhu, Zhuoyuan, Han, Dengke, Wang, Duo, Li, Wenming, Ye, Xiaochun, Hu, Cunchen, Chen, Hongyang, Sun, Guangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
von: Han, Dengke, et al.
Veröffentlicht: (2025)
von: Han, Dengke, et al.
Veröffentlicht: (2025)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Accelerating GNN Training through Locality-aware Dropout and Merge
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
von: Xue, Runzhen, et al.
Veröffentlicht: (2023)
von: Xue, Runzhen, et al.
Veröffentlicht: (2023)
Survey on Characterizing and Understanding GNNs from a Computer Architecture Perspective
von: Wu, Meng, et al.
Veröffentlicht: (2024)
von: Wu, Meng, et al.
Veröffentlicht: (2024)
ADE-HGNN: Accelerating HGNNs through Attention Disparity Exploitation
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
von: Wu, Meng, et al.
Veröffentlicht: (2024)
von: Wu, Meng, et al.
Veröffentlicht: (2024)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
von: Xue, Runzhen, et al.
Veröffentlicht: (2024)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space Exploration
von: Xue, Runzhen, et al.
Veröffentlicht: (2025)
von: Xue, Runzhen, et al.
Veröffentlicht: (2025)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
Control Flow Management in Modern GPUs
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
Topkima-Former: Low-energy, Low-Latency Inference for Transformers using top-k In-memory ADC
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
von: Dong, Shuai, et al.
Veröffentlicht: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
von: Kim, Hansung, et al.
Veröffentlicht: (2024)
von: Kim, Hansung, et al.
Veröffentlicht: (2024)
Mixed Structural Choice Operator: Enhancing Technology Mapping with Heterogeneous Representations
von: Hu, Zhang, et al.
Veröffentlicht: (2025)
von: Hu, Zhang, et al.
Veröffentlicht: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
CoverAssert: Iterative LLM Assertion Generation Driven by Functional Coverage via Syntax-Semantic Representations
von: Wang, Yonghao, et al.
Veröffentlicht: (2026)
von: Wang, Yonghao, et al.
Veröffentlicht: (2026)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event Cameras
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
AssertGen: Enhancement of LLM-aided Assertion Generation through Cross-Layer Signal Bridging
von: Lyu, Hongqin, et al.
Veröffentlicht: (2025)
von: Lyu, Hongqin, et al.
Veröffentlicht: (2025)
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
von: Hao, Mingbo, et al.
Veröffentlicht: (2026)
von: Hao, Mingbo, et al.
Veröffentlicht: (2026)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
von: Yu, Zhongkai, et al.
Veröffentlicht: (2024)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2024)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
von: Wang, Chuanning, et al.
Veröffentlicht: (2024)
von: Wang, Chuanning, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
von: Han, Dengke, et al.
Veröffentlicht: (2025) -
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024) -
Accelerating GNN Training through Locality-aware Dropout and Merge
von: Sun, Gongjian, et al.
Veröffentlicht: (2025) -
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
von: Xue, Runzhen, et al.
Veröffentlicht: (2023) -
Survey on Characterizing and Understanding GNNs from a Computer Architecture Perspective
von: Wu, Meng, et al.
Veröffentlicht: (2024)