Faster Inference of LLMs using FP8 on the Intel Gaudi
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Joonhyung, Markovich-Golan, Shmulik, Ohayon, Daniel, Hanani, Yair, Park, Gunho, Kim, Byeongwook, Karnieli, Asaf, Livne, Uri, Shen, Haihao, Huang, Tai, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Automatic mixed precision for optimizing gained time with constrained loss mean-squared-error based on model partition to sequential sub-graphs
by: Markovich-Golan, Shmulik, et al.
Published: (2025)
by: Markovich-Golan, Shmulik, et al.
Published: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
by: Gafni, Tomer, et al.
Published: (2025)
by: Gafni, Tomer, et al.
Published: (2025)
GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors
by: Zhang, Chengming, et al.
Published: (2024)
by: Zhang, Chengming, et al.
Published: (2024)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
by: Zou, Jiaxiang, et al.
Published: (2026)
by: Zou, Jiaxiang, et al.
Published: (2026)
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)
by: Wu, Hui, et al.
Published: (2023)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
by: Mukunoki, Daichi
Published: (2025)
by: Mukunoki, Daichi
Published: (2025)
if-ZKP: Intel FPGA-Based Acceleration of Zero Knowledge Proofs
by: Butt, Shahzad Ahmad, et al.
Published: (2024)
by: Butt, Shahzad Ahmad, et al.
Published: (2024)
NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16
by: Schmulbach, Viansa, et al.
Published: (2025)
by: Schmulbach, Viansa, et al.
Published: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)
by: Park, Gunho, et al.
Published: (2022)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
by: Gowda, Bindu G, et al.
Published: (2025)
by: Gowda, Bindu G, et al.
Published: (2025)
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
by: Kuper, Reese, et al.
Published: (2023)
by: Kuper, Reese, et al.
Published: (2023)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Tensor-Parallelism with Partially Synchronized Activations
by: Lamprecht, Itay, et al.
Published: (2025)
by: Lamprecht, Itay, et al.
Published: (2025)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
by: Oh, Dongsuk, et al.
Published: (2025)
by: Oh, Dongsuk, et al.
Published: (2025)
From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
by: Kwon, Omin, et al.
Published: (2025)
by: Kwon, Omin, et al.
Published: (2025)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN Performance
by: Kang, Seungkwan, et al.
Published: (2026)
by: Kang, Seungkwan, et al.
Published: (2026)
THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX
by: Dizani, Farshad, et al.
Published: (2025)
by: Dizani, Farshad, et al.
Published: (2025)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
by: Jang, Yongjoo, et al.
Published: (2025)
by: Jang, Yongjoo, et al.
Published: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
by: Nikolić, Miloš, et al.
Published: (2022)
by: Nikolić, Miloš, et al.
Published: (2022)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
by: Kwon, Jaewon, et al.
Published: (2025)
by: Kwon, Jaewon, et al.
Published: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
by: Mamdouh, Ahmed, et al.
Published: (2024)
by: Mamdouh, Ahmed, et al.
Published: (2024)
Bancroft: Genomics Acceleration Beyond On-Device Memory
by: Lim, Se-Min, et al.
Published: (2025)
by: Lim, Se-Min, et al.
Published: (2025)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
Accelerating AI and Computer Vision for Satellite Pose Estimation on the Intel Myriad X Embedded SoC
by: Leon, Vasileios, et al.
Published: (2024)
by: Leon, Vasileios, et al.
Published: (2024)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
by: Yoon, Dongho, et al.
Published: (2025)
by: Yoon, Dongho, et al.
Published: (2025)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
Similar Items
-
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025) -
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025) -
Automatic mixed precision for optimizing gained time with constrained loss mean-squared-error based on model partition to sequential sub-graphs
by: Markovich-Golan, Shmulik, et al.
Published: (2025) -
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024) -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)