Metis: Training LLMs with FP4 Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Hengjie, Chen, Mengyi, Yang, Yifeng, Huang, Ruijun, Dong, Fang, Zhou, Jixian, Chen, Anrui, Dong, Mingzhi, Wang, Yujiang, Hou, Jinlong, Cheng, Yuan, Wu, Fan, Yang, Fan, Lu, Tun, Gu, Ning, Shang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Dispelling the Curse of Singularities in Neural Network Optimizations
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026)
by: Huang, Zhendong, et al.
Published: (2026)
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
by: Chen, Anrui, et al.
Published: (2026)
by: Chen, Anrui, et al.
Published: (2026)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
by: Huang, Ruijun, et al.
Published: (2026)
by: Huang, Ruijun, et al.
Published: (2026)
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
by: Shi, Yubin, et al.
Published: (2024)
by: Shi, Yubin, et al.
Published: (2024)
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
Feature-Indexed Federated Recommendation with Residual-Quantized Codebooks
by: Han, Mingzhe, et al.
Published: (2026)
by: Han, Mingzhe, et al.
Published: (2026)
Medical records condensation: a roadmap towards healthcare data democratisation
by: Wang, Yujiang, et al.
Published: (2023)
by: Wang, Yujiang, et al.
Published: (2023)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026)
by: Zhou, Huilin, et al.
Published: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
SliderQuant: Accurate Post-Training Quantization for LLMs
by: Wang, Shigeng, et al.
Published: (2026)
by: Wang, Shigeng, et al.
Published: (2026)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
Assessment of MXD3 Expression as a Predictor of Survival in Lung Squamous Cell Carcinoma
by: Mingzhi Cao, et al.
Published: (2025)
by: Mingzhi Cao, et al.
Published: (2025)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
Enabling and Enacting: Interplay of Agency and Emotions in the Learning Experiences of Successful EFL Learners
by: Hengjie Chen, et al.
Published: (2025)
by: Hengjie Chen, et al.
Published: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
by: Liang, Guang, et al.
Published: (2025)
by: Liang, Guang, et al.
Published: (2025)
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers
by: Chen, Ruichen, et al.
Published: (2025)
by: Chen, Ruichen, et al.
Published: (2025)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
Biosensing materials for monoclonal antibody detection: An overview
by: Yuxin Zhang, et al.
Published: (2024)
by: Yuxin Zhang, et al.
Published: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
by: Shen, Leqi, et al.
Published: (2025)
by: Shen, Leqi, et al.
Published: (2025)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
by: Wei, Qikai, et al.
Published: (2024)
by: Wei, Qikai, et al.
Published: (2024)
Meerkat: A Distributed Reactive Programming Language with Live Updates
by: Zhong, Heng, et al.
Published: (2024)
by: Zhong, Heng, et al.
Published: (2024)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
FP3: A 3D Foundation Policy for Robotic Manipulation
by: Yang, Rujia, et al.
Published: (2025)
by: Yang, Rujia, et al.
Published: (2025)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
by: Lan, Xiaohan, et al.
Published: (2025)
by: Lan, Xiaohan, et al.
Published: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025)
by: Chen, Kun, et al.
Published: (2025)
QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
by: Li, Zhikai, et al.
Published: (2023)
by: Li, Zhikai, et al.
Published: (2023)
The Fort McKay Métis Nation
by: Fortna, Peter
Published: (2025)
by: Fortna, Peter
Published: (2025)
Mētis and violence in Machiavellian political theory
by: Regina Queiroz
Published: (2017)
by: Regina Queiroz
Published: (2017)
Bayesian experimental design: grouped geometric pooled posterior via ensemble Kalman methods
by: Yang, Huchen, et al.
Published: (2026)
by: Yang, Huchen, et al.
Published: (2026)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
by: Luo, Yingsong, et al.
Published: (2024)
by: Luo, Yingsong, et al.
Published: (2024)
InfiR2: A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models
by: Wang, Wenjun, et al.
Published: (2025)
by: Wang, Wenjun, et al.
Published: (2025)
Theory and Field to Determine the Construction of Bridge and Tunnel Linked Segment
by: Anrui Zhang, et al.
Published: (2026)
by: Anrui Zhang, et al.
Published: (2026)
Quantized nonlinear transport and its breakdown in Fermi gases with Berry curvature
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
Similar Items
-
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026) -
Dispelling the Curse of Singularities in Neural Network Optimizations
by: Cao, Hengjie, et al.
Published: (2026) -
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
by: Huang, Zhendong, et al.
Published: (2026) -
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
by: Chen, Anrui, et al.
Published: (2026) -
SD-MoE: Spectral Decomposition for Effective Expert Specialization
by: Huang, Ruijun, et al.
Published: (2026)