To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Joonhyung, Bae, Jeongin, Kim, Byeongwook, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
by: Bae, Jeongin, et al.
Published: (2026)
by: Bae, Jeongin, et al.
Published: (2026)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
by: Woo, Sunghyeon, et al.
Published: (2024)
by: Woo, Sunghyeon, et al.
Published: (2024)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
by: Kwon, Sangwoo, et al.
Published: (2025)
by: Kwon, Sangwoo, et al.
Published: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
by: Xi, Haocheng, et al.
Published: (2024)
by: Xi, Haocheng, et al.
Published: (2024)
Defeating the Training-Inference Mismatch via FP16
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
by: Lee, Yousung, et al.
Published: (2026)
by: Lee, Yousung, et al.
Published: (2026)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
by: Wang, Fengjuan, et al.
Published: (2025)
by: Wang, Fengjuan, et al.
Published: (2025)
Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models
by: Park, Youngrok, et al.
Published: (2025)
by: Park, Youngrok, et al.
Published: (2025)
OPC: One-Point-Contraction Unlearning Toward Deep Feature Forgetting
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
by: Lee, Dongseok, et al.
Published: (2025)
by: Lee, Dongseok, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
DAFA: Distance-Aware Fair Adversarial Training
by: Lee, Hyungyu, et al.
Published: (2024)
by: Lee, Hyungyu, et al.
Published: (2024)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Sparse Autoencoders, Again?
by: Lu, Yin, et al.
Published: (2025)
by: Lu, Yin, et al.
Published: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
by: Liu, Licheng, et al.
Published: (2025)
by: Liu, Licheng, et al.
Published: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
by: Zhang, Wuyue, et al.
Published: (2026)
by: Zhang, Wuyue, et al.
Published: (2026)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)
by: Kim, Hoyong, et al.
Published: (2023)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
Born Again Neural Networks
by: Furlanello, Tommaso, et al.
Published: (2018)
by: Furlanello, Tommaso, et al.
Published: (2018)
Provably Robust Pre-Trained Ensembles for Biomarker-Based Cancer Classification
by: Lee, Chongmin, et al.
Published: (2024)
by: Lee, Chongmin, et al.
Published: (2024)
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
by: Park, Taekhyun, et al.
Published: (2026)
by: Park, Taekhyun, et al.
Published: (2026)
VarDrop: Enhancing Training Efficiency by Reducing Variate Redundancy in Periodic Time Series Forecasting
by: Kang, Junhyeok, et al.
Published: (2025)
by: Kang, Junhyeok, et al.
Published: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
Similar Items
-
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025) -
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025) -
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024) -
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025) -
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)