$μ$nit Scaling: Simple and Scalable FP8 LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Narayan, Saaketh, Gupta, Abhay, Paul, Mansheej, Blalock, Davis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlashOptim: Optimizers for Memory-Efficient Training
von: Ortiz, Jose Javier Gonzalez, et al.
Veröffentlicht: (2026)
von: Ortiz, Jose Javier Gonzalez, et al.
Veröffentlicht: (2026)
Towards Fully FP8 GEMM LLM Training at Scale
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
Scaling FP8 training to trillion-token LLMs
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
Metis: Training LLMs with FP4 Quantization
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
Critique-out-Loud Reward Models
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
Boltzmann Reinforcement Learning for Noise resilience in Analog Ising Machines
von: Choudhary, Aditya, et al.
Veröffentlicht: (2026)
von: Choudhary, Aditya, et al.
Veröffentlicht: (2026)
Scaling Laws for Precision
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Elucidating the Design Space of FP4 training
von: Hu, Robert, et al.
Veröffentlicht: (2025)
von: Hu, Robert, et al.
Veröffentlicht: (2025)
Soup to go: mitigating forgetting during continual learning with model averaging
von: Kleiman, Anat, et al.
Veröffentlicht: (2025)
von: Kleiman, Anat, et al.
Veröffentlicht: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
von: Blakeney, Cody, et al.
Veröffentlicht: (2024)
von: Blakeney, Cody, et al.
Veröffentlicht: (2024)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
LLM-AutoSciLab: Closed-Loop Scientific Discovery via Active Experimentation with LLMs
von: Kabra, Sanchit, et al.
Veröffentlicht: (2026)
von: Kabra, Sanchit, et al.
Veröffentlicht: (2026)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
von: Baek, Daehyeon, et al.
Veröffentlicht: (2025)
von: Baek, Daehyeon, et al.
Veröffentlicht: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
von: Castro, Roberto L., et al.
Veröffentlicht: (2025)
von: Castro, Roberto L., et al.
Veröffentlicht: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
von: Liang, Guang, et al.
Veröffentlicht: (2025)
von: Liang, Guang, et al.
Veröffentlicht: (2025)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
von: Ankner, Zachary, et al.
Veröffentlicht: (2024)
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
von: Rajabi, Sahar, et al.
Veröffentlicht: (2025)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
von: Thangarasa, Vithursan, et al.
Veröffentlicht: (2023)
von: Thangarasa, Vithursan, et al.
Veröffentlicht: (2023)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
von: Nikolić, Miloš, et al.
Veröffentlicht: (2022)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
von: Li, Yitong, et al.
Veröffentlicht: (2026)
von: Li, Yitong, et al.
Veröffentlicht: (2026)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
von: Scetbon, Meyer, et al.
Veröffentlicht: (2025)
von: Scetbon, Meyer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlashOptim: Optimizers for Memory-Efficient Training
von: Ortiz, Jose Javier Gonzalez, et al.
Veröffentlicht: (2026) -
Towards Fully FP8 GEMM LLM Training at Scale
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025) -
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025) -
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024) -
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)