Guardado en:
| Autores principales: | Narayan, Saaketh, Gupta, Abhay, Paul, Mansheej, Blalock, Davis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.05967 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FlashOptim: Optimizers for Memory-Efficient Training
por: Ortiz, Jose Javier Gonzalez, et al.
Publicado: (2026)
por: Ortiz, Jose Javier Gonzalez, et al.
Publicado: (2026)
Towards Fully FP8 GEMM LLM Training at Scale
por: Hernández-Cano, Alejandro, et al.
Publicado: (2025)
por: Hernández-Cano, Alejandro, et al.
Publicado: (2025)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
por: Lee, Joonhyung, et al.
Publicado: (2024)
por: Lee, Joonhyung, et al.
Publicado: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
por: Kim, Jiwoo, et al.
Publicado: (2025)
por: Kim, Jiwoo, et al.
Publicado: (2025)
Scaling FP8 training to trillion-token LLMs
por: Fishman, Maxim, et al.
Publicado: (2024)
por: Fishman, Maxim, et al.
Publicado: (2024)
Critique-out-Loud Reward Models
por: Ankner, Zachary, et al.
Publicado: (2024)
por: Ankner, Zachary, et al.
Publicado: (2024)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
por: Cao, Hengjie, et al.
Publicado: (2026)
por: Cao, Hengjie, et al.
Publicado: (2026)
FP8 Quantization: The Power of the Exponent
por: Kuzmin, Andrey, et al.
Publicado: (2022)
por: Kuzmin, Andrey, et al.
Publicado: (2022)
Scaling Laws for Precision
por: Kumar, Tanishq, et al.
Publicado: (2024)
por: Kumar, Tanishq, et al.
Publicado: (2024)
Boltzmann Reinforcement Learning for Noise resilience in Analog Ising Machines
por: Choudhary, Aditya, et al.
Publicado: (2026)
por: Choudhary, Aditya, et al.
Publicado: (2026)
Soup to go: mitigating forgetting during continual learning with model averaging
por: Kleiman, Anat, et al.
Publicado: (2025)
por: Kleiman, Anat, et al.
Publicado: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
por: Blakeney, Cody, et al.
Publicado: (2024)
por: Blakeney, Cody, et al.
Publicado: (2024)
Metis: Training LLMs with FP4 Quantization
por: Cao, Hengjie, et al.
Publicado: (2025)
por: Cao, Hengjie, et al.
Publicado: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
por: Xi, Haocheng, et al.
Publicado: (2024)
por: Xi, Haocheng, et al.
Publicado: (2024)
LLM-AutoSciLab: Closed-Loop Scientific Discovery via Active Experimentation with LLMs
por: Kabra, Sanchit, et al.
Publicado: (2026)
por: Kabra, Sanchit, et al.
Publicado: (2026)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
por: Fujii, Kazuki, et al.
Publicado: (2024)
por: Fujii, Kazuki, et al.
Publicado: (2024)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
por: Zhang, Wuyue, et al.
Publicado: (2026)
por: Zhang, Wuyue, et al.
Publicado: (2026)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
por: Wang, Fengjuan, et al.
Publicado: (2025)
por: Wang, Fengjuan, et al.
Publicado: (2025)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
por: Kumar, Abhay, et al.
Publicado: (2025)
por: Kumar, Abhay, et al.
Publicado: (2025)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
por: Ankner, Zachary, et al.
Publicado: (2024)
por: Ankner, Zachary, et al.
Publicado: (2024)
Elucidating the Design Space of FP4 training
por: Hu, Robert, et al.
Publicado: (2025)
por: Hu, Robert, et al.
Publicado: (2025)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
por: Qiu, Zhaopeng, et al.
Publicado: (2026)
por: Qiu, Zhaopeng, et al.
Publicado: (2026)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
por: Xi, Haocheng, et al.
Publicado: (2026)
por: Xi, Haocheng, et al.
Publicado: (2026)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
por: Thangarasa, Vithursan, et al.
Publicado: (2023)
por: Thangarasa, Vithursan, et al.
Publicado: (2023)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
por: Baek, Daehyeon, et al.
Publicado: (2025)
por: Baek, Daehyeon, et al.
Publicado: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
por: Liang, Guang, et al.
Publicado: (2025)
por: Liang, Guang, et al.
Publicado: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
por: Chmiel, Brian, et al.
Publicado: (2025)
por: Chmiel, Brian, et al.
Publicado: (2025)
Defeating the Training-Inference Mismatch via FP16
por: Qi, Penghui, et al.
Publicado: (2025)
por: Qi, Penghui, et al.
Publicado: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
por: Castro, Roberto L., et al.
Publicado: (2025)
por: Castro, Roberto L., et al.
Publicado: (2025)
Muon is Scalable for LLM Training
por: Liu, Jingyuan, et al.
Publicado: (2025)
por: Liu, Jingyuan, et al.
Publicado: (2025)
Efficient Post-training Quantization with FP8 Formats
por: Shen, Haihao, et al.
Publicado: (2023)
por: Shen, Haihao, et al.
Publicado: (2023)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
por: Brinkmann, Jannik, et al.
Publicado: (2024)
por: Brinkmann, Jannik, et al.
Publicado: (2024)
Optimizing Large Language Model Training Using FP4 Quantization
por: Wang, Ruizhe, et al.
Publicado: (2025)
por: Wang, Ruizhe, et al.
Publicado: (2025)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
por: Li, Yitong, et al.
Publicado: (2026)
por: Li, Yitong, et al.
Publicado: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
por: Xia, Haojun, et al.
Publicado: (2024)
por: Xia, Haojun, et al.
Publicado: (2024)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
por: Nikolić, Miloš, et al.
Publicado: (2022)
por: Nikolić, Miloš, et al.
Publicado: (2022)
SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
por: Rajabi, Sahar, et al.
Publicado: (2025)
por: Rajabi, Sahar, et al.
Publicado: (2025)
WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
por: Wu, Junyu, et al.
Publicado: (2025)
por: Wu, Junyu, et al.
Publicado: (2025)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
Ejemplares similares
-
FlashOptim: Optimizers for Memory-Efficient Training
por: Ortiz, Jose Javier Gonzalez, et al.
Publicado: (2026) -
Towards Fully FP8 GEMM LLM Training at Scale
por: Hernández-Cano, Alejandro, et al.
Publicado: (2025) -
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
por: Zhang, Yu, et al.
Publicado: (2025) -
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
por: Lee, Joonhyung, et al.
Publicado: (2024) -
An Inquiry into Datacenter TCO for LLM Inference with FP8
por: Kim, Jiwoo, et al.
Publicado: (2025)