Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Binxing, Gu, Hao, Li, Lujun, Wang, Hao, Liu, Bei, Liu, Jiacheng, Zhu, Qiyuan, Yang, Xintong, Li, Chao, Han, Sirui, Guo, Yike |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
por: Yang, Xintong, et al.
Publicado: (2026)
por: Yang, Xintong, et al.
Publicado: (2026)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
por: Gu, Hao, et al.
Publicado: (2026)
por: Gu, Hao, et al.
Publicado: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
por: Maskey, Sohir, et al.
Publicado: (2026)
por: Maskey, Sohir, et al.
Publicado: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)
por: Zhang, Peiyuan, et al.
Publicado: (2026)
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
por: Dong, Peijie, et al.
Publicado: (2024)
por: Dong, Peijie, et al.
Publicado: (2024)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
por: Gernigon, Cédric, et al.
Publicado: (2024)
por: Gernigon, Cédric, et al.
Publicado: (2024)
SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
por: Li, Muyang, et al.
Publicado: (2024)
por: Li, Muyang, et al.
Publicado: (2024)
BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
por: Yan, Xiaobei, et al.
Publicado: (2025)
por: Yan, Xiaobei, et al.
Publicado: (2025)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
por: Li, Qingyuan, et al.
Publicado: (2025)
por: Li, Qingyuan, et al.
Publicado: (2025)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
por: Niu, Muqun, et al.
Publicado: (2024)
por: Niu, Muqun, et al.
Publicado: (2024)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
por: Liu, James, et al.
Publicado: (2024)
por: Liu, James, et al.
Publicado: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
por: Du, Dayou, et al.
Publicado: (2024)
por: Du, Dayou, et al.
Publicado: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Bit-Efficient Quantisation for Two-Channel Modulo-Sampling Systems
por: Yan, Wenyi, et al.
Publicado: (2026)
por: Yan, Wenyi, et al.
Publicado: (2026)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
Robust and Scalable Renaming with Subquadratic Bits
por: Bai, Sirui, et al.
Publicado: (2025)
por: Bai, Sirui, et al.
Publicado: (2025)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Rate-Splitting Multiple Access for Coexistence of Semantic and Bit Communications
por: Liu, Yuanwen, et al.
Publicado: (2024)
por: Liu, Yuanwen, et al.
Publicado: (2024)
BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
por: Aman, Euhid, et al.
Publicado: (2025)
por: Aman, Euhid, et al.
Publicado: (2025)
BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs
por: Karamat, Muhammad Zeeshan, et al.
Publicado: (2025)
por: Karamat, Muhammad Zeeshan, et al.
Publicado: (2025)
Bit by Bit: Gravity Through the Lens of Quantum Information
por: Munizzi, William
Publicado: (2024)
por: Munizzi, William
Publicado: (2024)
ResBit: Residual Bit Vector for Categorical Values
por: Fuchi, Masane, et al.
Publicado: (2023)
por: Fuchi, Masane, et al.
Publicado: (2023)
Delta Decompression for MoE-based LLMs Compression
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
Joint Power and Bit Allocation for Precoded Massive MIMO Channels
por: Liu, Shuiyin, et al.
Publicado: (2025)
por: Liu, Shuiyin, et al.
Publicado: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
ORBGRAND: Achievable Rate for General Bit Channels and Application in BICM
por: Li, Zhuang, et al.
Publicado: (2024)
por: Li, Zhuang, et al.
Publicado: (2024)
BitNet Distillation
por: Wu, Xun, et al.
Publicado: (2025)
por: Wu, Xun, et al.
Publicado: (2025)
Pixel-Bit
Publicado: (2021)
Publicado: (2021)
Bits and Pieces
por: O'Brien, Sarah
Publicado: (2023)
por: O'Brien, Sarah
Publicado: (2023)
A method of using RSVD in residual calculation of LowBit GEMM
por: Gu, Hongyaoxing
Publicado: (2024)
por: Gu, Hongyaoxing
Publicado: (2024)
From Bit to Block: Decoding on Erasure Channels
por: Pfister, Henry D., et al.
Publicado: (2025)
por: Pfister, Henry D., et al.
Publicado: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
Distributed Renaming with Subquadratic Bits via Scalable Committee Election
por: Bai, Sirui, et al.
Publicado: (2026)
por: Bai, Sirui, et al.
Publicado: (2026)
FrameQuant: Flexible Low-Bit Quantization for Transformers
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
Ejemplares similares
-
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
por: Gu, Hao, et al.
Publicado: (2025) -
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
por: Yang, Xintong, et al.
Publicado: (2026) -
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
por: Gu, Hao, et al.
Publicado: (2026) -
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025) -
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
por: Maskey, Sohir, et al.
Publicado: (2026)