I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Xing, Cheng, Yuan, Yang, Dawei, Yuan, Zhihang, Yu, Jiangyong, Xu, Chen, Zhou, Sifan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
por: Hu, Xing, et al.
Publicado: (2025)
por: Hu, Xing, et al.
Publicado: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025)
por: Xu, Zukang, et al.
Publicado: (2025)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
por: Hu, Xing, et al.
Publicado: (2025)
por: Hu, Xing, et al.
Publicado: (2025)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
por: Xu, Chen, et al.
Publicado: (2025)
por: Xu, Chen, et al.
Publicado: (2025)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
por: Yue, Yuxuan, et al.
Publicado: (2024)
por: Yue, Yuxuan, et al.
Publicado: (2024)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
por: Yu, Jiangyong, et al.
Publicado: (2025)
por: Yu, Jiangyong, et al.
Publicado: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
por: Yu, Jiangyong, et al.
Publicado: (2025)
por: Yu, Jiangyong, et al.
Publicado: (2025)
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization
por: Yu, JiangYong, et al.
Publicado: (2025)
por: Yu, JiangYong, et al.
Publicado: (2025)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
por: Xu, Zukang, et al.
Publicado: (2025)
por: Xu, Zukang, et al.
Publicado: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
por: Hu, Xing, et al.
Publicado: (2026)
por: Hu, Xing, et al.
Publicado: (2026)
DLLMQuant: Quantizing Diffusion-based Large Language Models
por: Xu, Chen, et al.
Publicado: (2025)
por: Xu, Chen, et al.
Publicado: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
por: Zhang, Chao, et al.
Publicado: (2025)
por: Zhang, Chao, et al.
Publicado: (2025)
Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
por: Liu, Wanlong, et al.
Publicado: (2025)
por: Liu, Wanlong, et al.
Publicado: (2025)
A Survey on Efficient Inference for Large Language Models
por: Zhou, Zixuan, et al.
Publicado: (2024)
por: Zhou, Zixuan, et al.
Publicado: (2024)
Information Entropy Guided Height-aware Histogram for Quantization-friendly Pillar Feature Encoder
por: Zhou, Sifan, et al.
Publicado: (2024)
por: Zhou, Sifan, et al.
Publicado: (2024)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
por: Duanmu, Haojie, et al.
Publicado: (2024)
por: Duanmu, Haojie, et al.
Publicado: (2024)
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
por: Yuan, Zhihang, et al.
Publicado: (2023)
por: Yuan, Zhihang, et al.
Publicado: (2023)
OneBit: Towards Extremely Low-bit Large Language Models
por: Xu, Yuzhuang, et al.
Publicado: (2024)
por: Xu, Yuzhuang, et al.
Publicado: (2024)
NLI:Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
por: Yu, Jiangyong, et al.
Publicado: (2026)
por: Yu, Jiangyong, et al.
Publicado: (2026)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
por: Wang, YiFeng, et al.
Publicado: (2026)
por: Wang, YiFeng, et al.
Publicado: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
por: Zhao, Pengxiang, et al.
Publicado: (2026)
por: Zhao, Pengxiang, et al.
Publicado: (2026)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
por: Zhang, Wenzheng, et al.
Publicado: (2026)
por: Zhang, Wenzheng, et al.
Publicado: (2026)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
por: Yue, Yuxuan, et al.
Publicado: (2025)
por: Yue, Yuxuan, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
por: Zeng, Chao, et al.
Publicado: (2024)
por: Zeng, Chao, et al.
Publicado: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
por: Liao, Baohao, et al.
Publicado: (2024)
por: Liao, Baohao, et al.
Publicado: (2024)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
por: Chee, Jerry, et al.
Publicado: (2023)
por: Chee, Jerry, et al.
Publicado: (2023)
Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
por: Yuan, Chenhan, et al.
Publicado: (2024)
por: Yuan, Chenhan, et al.
Publicado: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
por: Adepu, Harshavardhan, et al.
Publicado: (2024)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
por: Lv, Keyu, et al.
Publicado: (2026)
por: Lv, Keyu, et al.
Publicado: (2026)
Majority Bit-Aware Watermarking For Large Language Models
por: Xu, Jiahao, et al.
Publicado: (2025)
por: Xu, Jiahao, et al.
Publicado: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
por: Li, Zeyu, et al.
Publicado: (2025)
por: Li, Zeyu, et al.
Publicado: (2025)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
por: Lu, Yu-Chen, et al.
Publicado: (2025)
por: Lu, Yu-Chen, et al.
Publicado: (2025)
Ejemplares similares
-
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025) -
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
por: Hu, Xing, et al.
Publicado: (2025) -
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
por: Xu, Zukang, et al.
Publicado: (2025) -
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
por: Hu, Xing, et al.
Publicado: (2025) -
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
por: Xu, Chen, et al.
Publicado: (2025)