Efficient LLM inference solution on Intel GPU
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Hui, Gan, Yi, Yuan, Feng, Ma, Jing, Zhu, Wei, Xu, Yutao, Zhu, Hong, Zhu, Yuhua, Liu, Xiaoli, Gu, Jinghui, Zhao, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
by: Tyagi, Abhishek, et al.
Published: (2026)
by: Tyagi, Abhishek, et al.
Published: (2026)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
by: Zuo, Dongsheng, et al.
Published: (2025)
by: Zuo, Dongsheng, et al.
Published: (2025)
BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sampling
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
RoboGPU: Accelerating GPU Collision Detection for Robotics
by: Liu, Lufei, et al.
Published: (2026)
by: Liu, Lufei, et al.
Published: (2026)
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
by: Kuper, Reese, et al.
Published: (2023)
by: Kuper, Reese, et al.
Published: (2023)
Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2
by: Abreu, Steven, et al.
Published: (2025)
by: Abreu, Steven, et al.
Published: (2025)
NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16
by: Schmulbach, Viansa, et al.
Published: (2025)
by: Schmulbach, Viansa, et al.
Published: (2025)
Nebula: Enable City-Scale 3D Gaussian Splatting in Virtual Reality via Collaborative Rendering and Accelerated Stereo Rasterization
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
Technology solutions targeting the performance of gen-AI inference in resource constrained platforms
by: Kundu, Joyjit, et al.
Published: (2026)
by: Kundu, Joyjit, et al.
Published: (2026)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
by: Tan, Zhanhong, et al.
Published: (2024)
by: Tan, Zhanhong, et al.
Published: (2024)
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026)
by: Bai, Kangbo, et al.
Published: (2026)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025)
by: Zhu, Tianhao, et al.
Published: (2025)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
by: Cai, Jingwei, et al.
Published: (2025)
by: Cai, Jingwei, et al.
Published: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
by: Huang, Xiaotong, et al.
Published: (2025)
by: Huang, Xiaotong, et al.
Published: (2025)
if-ZKP: Intel FPGA-Based Acceleration of Zero Knowledge Proofs
by: Butt, Shahzad Ahmad, et al.
Published: (2024)
by: Butt, Shahzad Ahmad, et al.
Published: (2024)
CuLifter: Lifting GPU Binaries to Typed IR
by: Zhao, Jisheng, et al.
Published: (2026)
by: Zhao, Jisheng, et al.
Published: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
by: Zhao, Chunyuan, et al.
Published: (2025)
by: Zhao, Chunyuan, et al.
Published: (2025)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
by: Li, Hongyi, et al.
Published: (2025)
by: Li, Hongyi, et al.
Published: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
by: Zhu, Wenbin, et al.
Published: (2025)
by: Zhu, Wenbin, et al.
Published: (2025)
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
by: Zhou, Yaoyun, et al.
Published: (2025)
by: Zhou, Yaoyun, et al.
Published: (2025)
From Indiscriminate to Targeted: Efficient RTL Verification via Functionally Key Signal-Driven LLM Assertion Generation
by: Wang, Yonghao, et al.
Published: (2026)
by: Wang, Yonghao, et al.
Published: (2026)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
by: Gu, Yufeng, et al.
Published: (2025)
by: Gu, Yufeng, et al.
Published: (2025)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
by: Tang, Jiaping, et al.
Published: (2025)
by: Tang, Jiaping, et al.
Published: (2025)
EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
by: Huang, Yi, et al.
Published: (2025)
by: Huang, Yi, et al.
Published: (2025)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
Design of a GPU with Heterogeneous Cores for Graphics
by: Tomás, Aurora, et al.
Published: (2026)
by: Tomás, Aurora, et al.
Published: (2026)
Similar Items
-
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
by: Tyagi, Abhishek, et al.
Published: (2026) -
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024) -
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025) -
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
by: He, Siyuan, et al.
Published: (2025) -
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
by: Zuo, Dongsheng, et al.
Published: (2025)