BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ji-Fu, Zhang, Manyi, Xia, Xiaobo, Bao, Han, Bai, Haoli, Dong, Zhenhua, Yu, Xianzhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QuantClaw: Precision Where It Matters for OpenClaw
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
by: Lv, Keyu, et al.
Published: (2026)
by: Lv, Keyu, et al.
Published: (2026)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025)
by: Liu, Ruikang, et al.
Published: (2025)
Block Rotation is All You Need for MXFP4 Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
FreeAct: Freeing Activations for LLM Quantization
by: Liu, Xiaohao, et al.
Published: (2026)
by: Liu, Xiaohao, et al.
Published: (2026)
FlatQuant: Flatness Matters for LLM Quantization
by: Sun, Yuxuan, et al.
Published: (2024)
by: Sun, Yuxuan, et al.
Published: (2024)
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization
by: Fu, Zhongqian, et al.
Published: (2025)
by: Fu, Zhongqian, et al.
Published: (2025)
Faster and Better LLMs via Latency-Aware Test-Time Scaling
by: Wang, Zili, et al.
Published: (2025)
by: Wang, Zili, et al.
Published: (2025)
L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
by: Yuan, Tao, et al.
Published: (2025)
by: Yuan, Tao, et al.
Published: (2025)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
by: Jang, Wonsuk, et al.
Published: (2025)
by: Jang, Wonsuk, et al.
Published: (2025)
Accurate KV Cache Quantization with Outlier Tokens Tracing
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
by: Zhang, Hengyuan, et al.
Published: (2026)
by: Zhang, Hengyuan, et al.
Published: (2026)
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
by: Liu, Ruikang, et al.
Published: (2024)
by: Liu, Ruikang, et al.
Published: (2024)
A Simple Linear Patch Revives Layer-Pruned Large Language Models
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
by: Lin, Haokun, et al.
Published: (2024)
by: Lin, Haokun, et al.
Published: (2024)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
by: Wang, Yun, et al.
Published: (2026)
by: Wang, Yun, et al.
Published: (2026)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
by: Tian, Anxin, et al.
Published: (2026)
by: Tian, Anxin, et al.
Published: (2026)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
dMoE: dLLMs with Learnable Block Experts
by: Feng, Sicheng, et al.
Published: (2026)
by: Feng, Sicheng, et al.
Published: (2026)
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
by: Leng, Jixuan, et al.
Published: (2026)
by: Leng, Jixuan, et al.
Published: (2026)
Training LLMs with MXFP4
by: Tseng, Albert, et al.
Published: (2025)
by: Tseng, Albert, et al.
Published: (2025)
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
by: Cim, Musa, et al.
Published: (2026)
by: Cim, Musa, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
OutlierTune: Efficient Channel-Wise Quantization for Large Language Models
by: Wang, Jinguang, et al.
Published: (2024)
by: Wang, Jinguang, et al.
Published: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
by: Gordon, Ofir, et al.
Published: (2026)
by: Gordon, Ofir, et al.
Published: (2026)
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
by: Xu, Bingxin, et al.
Published: (2025)
by: Xu, Bingxin, et al.
Published: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
by: Yao, Dingyu, et al.
Published: (2025)
by: Yao, Dingyu, et al.
Published: (2025)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
by: Bao, Wenrui, et al.
Published: (2025)
by: Bao, Wenrui, et al.
Published: (2025)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
by: Chen, Jierun, et al.
Published: (2025)
by: Chen, Jierun, et al.
Published: (2025)
Accurate Block Quantization in LLMs with Outliers
by: Trukhanov, Nikita, et al.
Published: (2024)
by: Trukhanov, Nikita, et al.
Published: (2024)
Similar Items
-
QuantClaw: Precision Where It Matters for OpenClaw
by: Zhang, Manyi, et al.
Published: (2026) -
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026) -
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
by: Lv, Keyu, et al.
Published: (2026) -
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025) -
Block Rotation is All You Need for MXFP4 Quantization
by: Shao, Yuantian, et al.
Published: (2025)