FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Fengjuan, Su, Zhiyi, Hu, Xingzhu, Wang, Cheng, Sun, Mou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
Scaling FP8 training to trillion-token LLMs
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
InfiR2: A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models
von: Wang, Wenjun, et al.
Veröffentlicht: (2025)
von: Wang, Wenjun, et al.
Veröffentlicht: (2025)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2026)
von: Xu, Zukang, et al.
Veröffentlicht: (2026)
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
von: Wu, Haoze, et al.
Veröffentlicht: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
von: Duanmu, Haojie, et al.
Veröffentlicht: (2025)
von: Duanmu, Haojie, et al.
Veröffentlicht: (2025)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
von: Hu, Gang, et al.
Veröffentlicht: (2025)
von: Hu, Gang, et al.
Veröffentlicht: (2025)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
von: Zhao, Jiayu, et al.
Veröffentlicht: (2026)
Auto-FP: An Experimental Study of Automated Feature Preprocessing for Tabular Data
von: Qi, Danrui, et al.
Veröffentlicht: (2023)
von: Qi, Danrui, et al.
Veröffentlicht: (2023)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
von: Metere, Alfredo
Veröffentlicht: (2025)
von: Metere, Alfredo
Veröffentlicht: (2025)
GRIN: GRadient-INformed MoE
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
von: Liu, Liyuan, et al.
Veröffentlicht: (2024)
StreamFP: Learnable Fingerprint-guided Data Selection for Efficient Stream Learning
von: Shi, Tongjun, et al.
Veröffentlicht: (2024)
von: Shi, Tongjun, et al.
Veröffentlicht: (2024)
Pretraining large language models with MXFP4 on Native FP4 Hardware
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
von: Su, Yang, et al.
Veröffentlicht: (2025)
von: Su, Yang, et al.
Veröffentlicht: (2025)
Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
von: Zhong, Zhengjia, et al.
Veröffentlicht: (2026)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026) -
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023) -
Scaling FP8 training to trillion-token LLMs
von: Fishman, Maxim, et al.
Veröffentlicht: (2024) -
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026) -
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)