ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Meng, Haoqian, Luo, Yilun, Zhao, Yafei, Liu, Wenyuan, Zhang, Peng, Ma, Xindian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
di: Luo, Yilun, et al.
Pubblicazione: (2025)
di: Luo, Yilun, et al.
Pubblicazione: (2025)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
di: Liu, Wenyuan, et al.
Pubblicazione: (2024)
di: Liu, Wenyuan, et al.
Pubblicazione: (2024)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
di: Ma, Xindian, et al.
Pubblicazione: (2026)
di: Ma, Xindian, et al.
Pubblicazione: (2026)
FAAR: Format-Aware Adaptive Rounding for NVFP4
di: Li, Hanglin, et al.
Pubblicazione: (2026)
di: Li, Hanglin, et al.
Pubblicazione: (2026)
Pretraining Large Language Models with NVFP4
di: NVIDIA, et al.
Pubblicazione: (2025)
di: NVIDIA, et al.
Pubblicazione: (2025)
TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
di: Tao, Qian, et al.
Pubblicazione: (2024)
di: Tao, Qian, et al.
Pubblicazione: (2024)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
di: Xia, Haojun, et al.
Pubblicazione: (2025)
di: Xia, Haojun, et al.
Pubblicazione: (2025)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
di: Xiang, Maoyang, et al.
Pubblicazione: (2026)
di: Xiang, Maoyang, et al.
Pubblicazione: (2026)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
di: Zhang, Jing, et al.
Pubblicazione: (2024)
di: Zhang, Jing, et al.
Pubblicazione: (2024)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
di: Luo, Yingsong, et al.
Pubblicazione: (2024)
di: Luo, Yingsong, et al.
Pubblicazione: (2024)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2026)
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2026)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
di: Chmiel, Brian, et al.
Pubblicazione: (2025)
di: Chmiel, Brian, et al.
Pubblicazione: (2025)
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
Cognitive Load-Aware Inference: A Neuro-Symbolic Framework for Optimizing the Token Economy of Large Language Models
di: Zhang, Yilun
Pubblicazione: (2025)
di: Zhang, Yilun
Pubblicazione: (2025)
Support Vector Boosting Machine (SVBM): Enhancing Classification Performance with AdaBoost and Residual Connections
di: Lian, Junbo Jacob
Pubblicazione: (2024)
di: Lian, Junbo Jacob
Pubblicazione: (2024)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
di: Xin, Meng, et al.
Pubblicazione: (2026)
di: Xin, Meng, et al.
Pubblicazione: (2026)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
CSRA: Controlled Spectral Residual Augmentation for Robust Sepsis Prediction
di: Guo, Honglin, et al.
Pubblicazione: (2026)
di: Guo, Honglin, et al.
Pubblicazione: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
di: Zhang, Shenao, et al.
Pubblicazione: (2024)
di: Zhang, Shenao, et al.
Pubblicazione: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
di: Bao, Chengzhu, et al.
Pubblicazione: (2026)
di: Bao, Chengzhu, et al.
Pubblicazione: (2026)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
di: Zhao, Minda, et al.
Pubblicazione: (2026)
di: Zhao, Minda, et al.
Pubblicazione: (2026)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
Residual Cross-Attention Transformer-Based Multi-User CSI Feedback with Deep Joint Source-Channel Coding
di: Zhang, Hengwei, et al.
Pubblicazione: (2025)
di: Zhang, Hengwei, et al.
Pubblicazione: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
di: Li, Junlin, et al.
Pubblicazione: (2026)
di: Li, Junlin, et al.
Pubblicazione: (2026)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
Achieving binary weight and activation for LLMs using Post-Training Quantization
di: Song, Siqing, et al.
Pubblicazione: (2025)
di: Song, Siqing, et al.
Pubblicazione: (2025)
Automated Formalization via Conceptual Retrieval-Augmented LLMs
di: Lu, Wangyue, et al.
Pubblicazione: (2025)
di: Lu, Wangyue, et al.
Pubblicazione: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
di: Su, Le, et al.
Pubblicazione: (2026)
di: Su, Le, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
di: Liu, Wenyuan, et al.
Pubblicazione: (2025) -
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
di: Luo, Yilun, et al.
Pubblicazione: (2025) -
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
di: Liu, Wenyuan, et al.
Pubblicazione: (2024) -
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
di: Ma, Xindian, et al.
Pubblicazione: (2026) -
FAAR: Format-Aware Adaptive Rounding for NVFP4
di: Li, Hanglin, et al.
Pubblicazione: (2026)