An Inquiry into Datacenter TCO for LLM Inference with FP8
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Jiwoo, Lee, Joonhyung, Park, Gunho, Kim, Byeongwook, Kwon, Se Jung, Lee, Dongsoo, Lee, Youngjoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
di: Lee, Joonhyung, et al.
Pubblicazione: (2025)
di: Lee, Joonhyung, et al.
Pubblicazione: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
di: Yang, June Yong, et al.
Pubblicazione: (2024)
di: Yang, June Yong, et al.
Pubblicazione: (2024)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2026)
di: Chu, Kexin, et al.
Pubblicazione: (2026)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
Anatomizing Deep Learning Inference in Web Browsers
di: Wang, Qipeng, et al.
Pubblicazione: (2024)
di: Wang, Qipeng, et al.
Pubblicazione: (2024)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
di: Metere, Alfredo
Pubblicazione: (2025)
di: Metere, Alfredo
Pubblicazione: (2025)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
di: Wu, Wenhao, et al.
Pubblicazione: (2026)
di: Wu, Wenhao, et al.
Pubblicazione: (2026)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
di: Yang, Kichang, et al.
Pubblicazione: (2025)
di: Yang, Kichang, et al.
Pubblicazione: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
di: Ibrahim, Muhammad Sohail, et al.
Pubblicazione: (2024)
di: Ibrahim, Muhammad Sohail, et al.
Pubblicazione: (2024)
KVDirect: Distributed Disaggregated LLM Inference
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
di: Chen, Shiyang, et al.
Pubblicazione: (2024)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
di: Xue, Leyang, et al.
Pubblicazione: (2024)
di: Xue, Leyang, et al.
Pubblicazione: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
di: Holmes, Connor, et al.
Pubblicazione: (2024)
di: Holmes, Connor, et al.
Pubblicazione: (2024)
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
di: Sun, Wenbo, et al.
Pubblicazione: (2024)
di: Sun, Wenbo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024) -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025) -
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025) -
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025) -
Faster Inference of LLMs using FP8 on the Intel Gaudi
di: Lee, Joonhyung, et al.
Pubblicazione: (2025)