FrameQuant: Flexible Low-Bit Quantization for Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adepu, Harshavardhan, Zeng, Zhanpeng, Zhang, Li, Singh, Vikas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
von: Zeng, Zhanpeng, et al.
Veröffentlicht: (2024)
von: Zeng, Zhanpeng, et al.
Veröffentlicht: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
von: Vicentino, Caio
Veröffentlicht: (2026)
von: Vicentino, Caio
Veröffentlicht: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
von: Kim, Jinhee, et al.
Veröffentlicht: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
SpinQuant: LLM quantization with learned rotations
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
von: Battula, Harshavardhan, et al.
Veröffentlicht: (2024)
von: Battula, Harshavardhan, et al.
Veröffentlicht: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
LQER: Low-Rank Quantization Error Reconstruction for LLMs
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Multi-Bit Distortion-Free Watermarking for Large Language Models
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025)
von: Wu, Xun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IM-Unpack: Training and Inference with Arbitrarily Low Precision Integers
von: Zeng, Zhanpeng, et al.
Veröffentlicht: (2024) -
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026) -
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025) -
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026) -
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)