Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Georgiou, Athos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
von: Lo, Yun-Chen, et al.
Veröffentlicht: (2024)
von: Lo, Yun-Chen, et al.
Veröffentlicht: (2024)
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
Exploration of Cryptocurrency Mining-Specific GPUs in AI Applications: A Case Study of CMP 170HX
von: Kangwei, Xing
Veröffentlicht: (2025)
von: Kangwei, Xing
Veröffentlicht: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025)
von: Meng, William, et al.
Veröffentlicht: (2025)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
von: Makin, Yashasvi, et al.
Veröffentlicht: (2025)
von: Makin, Yashasvi, et al.
Veröffentlicht: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
von: Ortega, Cristobal, et al.
Veröffentlicht: (2024)
von: Ortega, Cristobal, et al.
Veröffentlicht: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
von: DeBole, Michael V., et al.
Veröffentlicht: (2025)
von: DeBole, Michael V., et al.
Veröffentlicht: (2025)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026) -
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026) -
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
von: Lo, Yun-Chen, et al.
Veröffentlicht: (2024) -
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
von: Antunes, Pedro, et al.
Veröffentlicht: (2025) -
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)