Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization
Fuente:
arXiv
Salvato in:
| Autori principali: | Xi, Haocheng, Chen, Yuxiang, Zhao, Kang, Teh, Kai Jun, Chen, Jianfei, Zhu, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accurate INT8 Training Through Dynamic Block-Level Fallback
di: Zhang, Pengle, et al.
Pubblicazione: (2025)
di: Zhang, Pengle, et al.
Pubblicazione: (2025)
Oscillation-Reduced MXFP4 Training for Vision Transformers
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
di: Zhang, Jintao, et al.
Pubblicazione: (2024)
di: Zhang, Jintao, et al.
Pubblicazione: (2024)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
di: Chen, Shimao, et al.
Pubblicazione: (2024)
di: Chen, Shimao, et al.
Pubblicazione: (2024)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
Accelerating Transformer Pre-training with 2:4 Sparsity
di: Hu, Yuezhou, et al.
Pubblicazione: (2024)
di: Hu, Yuezhou, et al.
Pubblicazione: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
di: Zhang, Jintao, et al.
Pubblicazione: (2024)
di: Zhang, Jintao, et al.
Pubblicazione: (2024)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
di: Xi, Haocheng, et al.
Pubblicazione: (2024)
di: Xi, Haocheng, et al.
Pubblicazione: (2024)
Efficient Backpropagation with Variance-Controlled Adaptive Sampling
di: Wang, Ziteng, et al.
Pubblicazione: (2024)
di: Wang, Ziteng, et al.
Pubblicazione: (2024)
S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training
di: Hu, Yuezhou, et al.
Pubblicazione: (2024)
di: Hu, Yuezhou, et al.
Pubblicazione: (2024)
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
di: Zhao, Yilong, et al.
Pubblicazione: (2023)
di: Zhao, Yilong, et al.
Pubblicazione: (2023)
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
di: Wang, Ziteng, et al.
Pubblicazione: (2024)
di: Wang, Ziteng, et al.
Pubblicazione: (2024)
Efficient Hyperparameter Tuning via Trajectory Invariance Principle
di: Li, Bingrui, et al.
Pubblicazione: (2025)
di: Li, Bingrui, et al.
Pubblicazione: (2025)
QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
di: Liu, Jing, et al.
Pubblicazione: (2023)
di: Liu, Jing, et al.
Pubblicazione: (2023)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
di: Baek, Daehyeon, et al.
Pubblicazione: (2025)
di: Baek, Daehyeon, et al.
Pubblicazione: (2025)
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
di: Ma, Jingxiao, et al.
Pubblicazione: (2025)
di: Ma, Jingxiao, et al.
Pubblicazione: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
SparseDM: Toward Sparse Efficient Diffusion Models
di: Wang, Kafeng, et al.
Pubblicazione: (2024)
di: Wang, Kafeng, et al.
Pubblicazione: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs
di: Zheng, Kaiwen, et al.
Pubblicazione: (2023)
di: Zheng, Kaiwen, et al.
Pubblicazione: (2023)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
di: Gao, Chang, et al.
Pubblicazione: (2024)
di: Gao, Chang, et al.
Pubblicazione: (2024)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024)
di: Li, Bingrui, et al.
Pubblicazione: (2024)
When Flat Minima Fail: Characterizing INT4 Quantization Collapse After FP32 Convergence
di: Armstrong, Marcus
Pubblicazione: (2026)
di: Armstrong, Marcus
Pubblicazione: (2026)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
di: Huang, Weiyu, et al.
Pubblicazione: (2025)
di: Huang, Weiyu, et al.
Pubblicazione: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
di: Jia, Jinda, et al.
Pubblicazione: (2026)
di: Jia, Jinda, et al.
Pubblicazione: (2026)
Diffusion Bridge Implicit Models
di: Zheng, Kaiwen, et al.
Pubblicazione: (2024)
di: Zheng, Kaiwen, et al.
Pubblicazione: (2024)
C-GAIL: Stabilizing Generative Adversarial Imitation Learning with Control Theory
di: Luo, Tianjiao, et al.
Pubblicazione: (2024)
di: Luo, Tianjiao, et al.
Pubblicazione: (2024)
SaberLDA: Sparsity-Aware Learning of Topic Models on GPUs
di: Li, Kaiwei, et al.
Pubblicazione: (2016)
di: Li, Kaiwei, et al.
Pubblicazione: (2016)
FlowDA: Accurate, Low-Latency Weather Data Assimilation via Flow Matching
di: Cheng, Ran, et al.
Pubblicazione: (2026)
di: Cheng, Ran, et al.
Pubblicazione: (2026)
Cross-Modal Reconstruction Pretraining for Ramp Flow Prediction at Highway Interchanges
di: Li, Yongchao, et al.
Pubblicazione: (2025)
di: Li, Yongchao, et al.
Pubblicazione: (2025)
Deterministic Differentiable Structured Pruning for Large Language Models
di: Huang, Weiyu, et al.
Pubblicazione: (2026)
di: Huang, Weiyu, et al.
Pubblicazione: (2026)
TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
di: Hwang, Younghye, et al.
Pubblicazione: (2025)
di: Hwang, Younghye, et al.
Pubblicazione: (2025)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
di: Kleinegger, Maximilian, et al.
Pubblicazione: (2026)
di: Kleinegger, Maximilian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Accurate INT8 Training Through Dynamic Block-Level Fallback
di: Zhang, Pengle, et al.
Pubblicazione: (2025) -
Oscillation-Reduced MXFP4 Training for Vision Transformers
di: Chen, Yuxiang, et al.
Pubblicazione: (2025) -
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
di: Zhang, Jintao, et al.
Pubblicazione: (2024) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
di: Chen, Shimao, et al.
Pubblicazione: (2024) -
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
di: Zhang, Jintao, et al.
Pubblicazione: (2025)