Pretraining large language models with MXFP4 on Native FP4 Hardware
Fuente:
arXiv
Saved in:
| Main Authors: | Cim, Musa, Palangappa, Poovaiah, Hodak, Miro, Dwivedula, Ravi, Arunachalam, Meena, Kandemir, Mahmut Taylan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
by: Cim, Musa, et al.
Published: (2026)
by: Cim, Musa, et al.
Published: (2026)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026)
by: Topcu, Burak, et al.
Published: (2026)
Parallel Context Compaction for Long-Horizon LLM Agent Serving
by: Cim, Musa, et al.
Published: (2026)
by: Cim, Musa, et al.
Published: (2026)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor
by: Li, Xiaocan, et al.
Published: (2026)
by: Li, Xiaocan, et al.
Published: (2026)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
Recipes for Pre-training LLMs with MXFP8
by: Mishra, Asit, et al.
Published: (2025)
by: Mishra, Asit, et al.
Published: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
by: Gulhan, Ahmed Burak, et al.
Published: (2025)
by: Gulhan, Ahmed Burak, et al.
Published: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
MXNorm: Reusing MXFP block scales for efficient tensor normalisation
by: McLean, Callum, et al.
Published: (2026)
by: McLean, Callum, et al.
Published: (2026)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference
by: Feng, Yilin, et al.
Published: (2026)
by: Feng, Yilin, et al.
Published: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
by: Zhang, Wuyue, et al.
Published: (2026)
by: Zhang, Wuyue, et al.
Published: (2026)
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks
by: Mei, Taiyuan, et al.
Published: (2024)
by: Mei, Taiyuan, et al.
Published: (2024)
Long-form factuality in large language models
by: Wei, Jerry, et al.
Published: (2024)
by: Wei, Jerry, et al.
Published: (2024)
Can large language models explore in-context?
by: Krishnamurthy, Akshay, et al.
Published: (2024)
by: Krishnamurthy, Akshay, et al.
Published: (2024)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
Early prediction of onset of sepsis in Clinical Setting
by: Mohammad, Fahim, et al.
Published: (2024)
by: Mohammad, Fahim, et al.
Published: (2024)
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
by: Zhou, Jiecheng, et al.
Published: (2025)
by: Zhou, Jiecheng, et al.
Published: (2025)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
by: Li, Ji-Fu, et al.
Published: (2026)
by: Li, Ji-Fu, et al.
Published: (2026)
A federated large language model for long-term time series forecasting
by: Abdel-Sater, Raed, et al.
Published: (2024)
by: Abdel-Sater, Raed, et al.
Published: (2024)
Are large language models superhuman chemists?
by: Mirza, Adrian, et al.
Published: (2024)
by: Mirza, Adrian, et al.
Published: (2024)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
by: Wang, Fengjuan, et al.
Published: (2025)
by: Wang, Fengjuan, et al.
Published: (2025)
Comprehensive benchmarking of large language models for RNA secondary structure prediction
by: Zablocki, L. I., et al.
Published: (2024)
by: Zablocki, L. I., et al.
Published: (2024)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
Using large language models for embodied planning introduces systematic safety risks
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
by: Zhang, Zhiyuan, et al.
Published: (2026)
by: Zhang, Zhiyuan, et al.
Published: (2026)
Similar Items
-
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
by: Cim, Musa, et al.
Published: (2026) -
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026) -
Parallel Context Compaction for Long-Horizon LLM Agent Serving
by: Cim, Musa, et al.
Published: (2026) -
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026) -
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
by: Xu, Zukang, et al.
Published: (2026)