Scaling FP8 training to trillion-token LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fishman, Maxim, Chmiel, Brian, Banner, Ron, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
Normalized Architectures are Natively 4-Bit
von: Fishman, Maxim, et al.
Veröffentlicht: (2026)
von: Fishman, Maxim, et al.
Veröffentlicht: (2026)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
Workspace Optimization: How to Train Your Agent
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
von: Chmiel, Brian, et al.
Veröffentlicht: (2021)
von: Chmiel, Brian, et al.
Veröffentlicht: (2021)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
Not all tokens are needed(NAT): token efficient reinforcement learning
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
Physics in Next-token Prediction
von: An, Hongjun, et al.
Veröffentlicht: (2024)
von: An, Hongjun, et al.
Veröffentlicht: (2024)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
von: Hoffer, Elad, et al.
Veröffentlicht: (2026)
von: Hoffer, Elad, et al.
Veröffentlicht: (2026)
Next-token pretraining implies in-context learning
von: Riechers, Paul M., et al.
Veröffentlicht: (2025)
von: Riechers, Paul M., et al.
Veröffentlicht: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
von: Maximov, Egor, et al.
Veröffentlicht: (2025)
von: Maximov, Egor, et al.
Veröffentlicht: (2025)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
Recipes for Pre-training LLMs with MXFP8
von: Mishra, Asit, et al.
Veröffentlicht: (2025)
von: Mishra, Asit, et al.
Veröffentlicht: (2025)
Bio2Token: All-atom tokenization of any biomolecular structure with Mamba
von: Liu, Andrew, et al.
Veröffentlicht: (2024)
von: Liu, Andrew, et al.
Veröffentlicht: (2024)
StreamFP: Learnable Fingerprint-guided Data Selection for Efficient Stream Learning
von: Shi, Tongjun, et al.
Veröffentlicht: (2024)
von: Shi, Tongjun, et al.
Veröffentlicht: (2024)
Pretraining large language models with MXFP4 on Native FP4 Hardware
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Scaling Laws for Pre-training Agents and World Models
von: Pearce, Tim, et al.
Veröffentlicht: (2024)
von: Pearce, Tim, et al.
Veröffentlicht: (2024)
OSF: On Pre-training and Scaling of Sleep Foundation Models
von: Shuai, Zitao, et al.
Veröffentlicht: (2026)
von: Shuai, Zitao, et al.
Veröffentlicht: (2026)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025) -
Normalized Architectures are Natively 4-Bit
von: Fishman, Maxim, et al.
Veröffentlicht: (2026) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022) -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024) -
Workspace Optimization: How to Train Your Agent
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)