When Flat Minima Fail: Characterizing INT4 Quantization Collapse After FP32 Convergence
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Armstrong, Marcus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
von: Baek, Daehyeon, et al.
Veröffentlicht: (2025)
von: Baek, Daehyeon, et al.
Veröffentlicht: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Are Flat Minima an Illusion?
von: Bennett, Michael Timothy
Veröffentlicht: (2026)
von: Bennett, Michael Timothy
Veröffentlicht: (2026)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
Metis: Training LLMs with FP4 Quantization
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Towards Robust Influence Functions with Flat Validation Minima
von: Ye, Xichen, et al.
Veröffentlicht: (2025)
von: Ye, Xichen, et al.
Veröffentlicht: (2025)
A PAC-Bayesian Link Between Generalisation and Flat Minima
von: Haddouche, Maxime, et al.
Veröffentlicht: (2024)
von: Haddouche, Maxime, et al.
Veröffentlicht: (2024)
Flat Minima and Generalization: Insights from Stochastic Convex Optimization
von: Schliserman, Matan, et al.
Veröffentlicht: (2025)
von: Schliserman, Matan, et al.
Veröffentlicht: (2025)
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
Towards the Connection between Activation Sparsity and Flat Minima
von: Peng, Ze, et al.
Veröffentlicht: (2026)
von: Peng, Ze, et al.
Veröffentlicht: (2026)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
Zeroth-Order Optimization Finds Flat Minima
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2025)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2025)
A Flat Minima Perspective on Understanding Augmentations and Model Robustness
von: Yoo, Weebum, et al.
Veröffentlicht: (2025)
von: Yoo, Weebum, et al.
Veröffentlicht: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
von: Chmiel, Brian, et al.
Veröffentlicht: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
When Adaptation Fails: A Gradient-Based Diagnosis of Collapsed Gating in Vision-Language Prompt Learning
von: Fang, Yunxuan, et al.
Veröffentlicht: (2026)
von: Fang, Yunxuan, et al.
Veröffentlicht: (2026)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
Jetfire: Efficient and Accurate Transformer Pretraining with INT8 Data Flow and Per-Block Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
A Function-Centric Perspective on Flat and Sharp Minima
von: Mason-Williams, Israel, et al.
Veröffentlicht: (2025)
von: Mason-Williams, Israel, et al.
Veröffentlicht: (2025)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
Flatness After All?
von: Shoham, Neta, et al.
Veröffentlicht: (2025)
von: Shoham, Neta, et al.
Veröffentlicht: (2025)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
von: Zhong, Shanshan, et al.
Veröffentlicht: (2024)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2024)
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
von: Han, Ting, et al.
Veröffentlicht: (2025)
von: Han, Ting, et al.
Veröffentlicht: (2025)
Noise Stability Optimization for Finding Flat Minima: A Hessian-based Regularization Approach
von: Zhang, Hongyang R., et al.
Veröffentlicht: (2023)
von: Zhang, Hongyang R., et al.
Veröffentlicht: (2023)
Theory-optimal Quantization Based on Flatness
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
Elucidating the Design Space of FP4 training
von: Hu, Robert, et al.
Veröffentlicht: (2025)
von: Hu, Robert, et al.
Veröffentlicht: (2025)
Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
von: Zheng, Meixi, et al.
Veröffentlicht: (2025)
von: Zheng, Meixi, et al.
Veröffentlicht: (2025)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
When Flatness Does (Not) Guarantee Adversarial Robustness
von: Walter, Nils Philipp, et al.
Veröffentlicht: (2025)
von: Walter, Nils Philipp, et al.
Veröffentlicht: (2025)
LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
von: Baek, Daehyeon, et al.
Veröffentlicht: (2025) -
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025) -
Are Flat Minima an Illusion?
von: Bennett, Michael Timothy
Veröffentlicht: (2026) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
von: Chen, Shimao, et al.
Veröffentlicht: (2024) -
Metis: Training LLMs with FP4 Quantization
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)