ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Meng, Haoqian, Luo, Yilun, Zhao, Yafei, Liu, Wenyuan, Zhang, Peng, Ma, Xindian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
por: Liu, Wenyuan, et al.
Publicado: (2025)
por: Liu, Wenyuan, et al.
Publicado: (2025)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
por: Luo, Yilun, et al.
Publicado: (2025)
por: Luo, Yilun, et al.
Publicado: (2025)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
por: Liu, Wenyuan, et al.
Publicado: (2024)
por: Liu, Wenyuan, et al.
Publicado: (2024)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
por: Ma, Xindian, et al.
Publicado: (2026)
por: Ma, Xindian, et al.
Publicado: (2026)
FAAR: Format-Aware Adaptive Rounding for NVFP4
por: Li, Hanglin, et al.
Publicado: (2026)
por: Li, Hanglin, et al.
Publicado: (2026)
Pretraining Large Language Models with NVFP4
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
por: Chen, Yuxiang, et al.
Publicado: (2025)
por: Chen, Yuxiang, et al.
Publicado: (2025)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
por: Tao, Qian, et al.
Publicado: (2024)
por: Tao, Qian, et al.
Publicado: (2024)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
por: Li, Qingyuan, et al.
Publicado: (2024)
por: Li, Qingyuan, et al.
Publicado: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
por: Liu, Yijun, et al.
Publicado: (2024)
por: Liu, Yijun, et al.
Publicado: (2024)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
por: Xia, Haojun, et al.
Publicado: (2025)
por: Xia, Haojun, et al.
Publicado: (2025)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
por: Xiang, Maoyang, et al.
Publicado: (2026)
por: Xiang, Maoyang, et al.
Publicado: (2026)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
por: Liu, Wanlong, et al.
Publicado: (2024)
por: Liu, Wanlong, et al.
Publicado: (2024)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
por: Zhang, Jing, et al.
Publicado: (2024)
por: Zhang, Jing, et al.
Publicado: (2024)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
por: Zhang, Zhiyuan, et al.
Publicado: (2026)
por: Zhang, Zhiyuan, et al.
Publicado: (2026)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
por: Zhao, Jiaqi, et al.
Publicado: (2025)
por: Zhao, Jiaqi, et al.
Publicado: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
por: Chmiel, Brian, et al.
Publicado: (2025)
por: Chmiel, Brian, et al.
Publicado: (2025)
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
por: Piękos, Piotr, et al.
Publicado: (2025)
por: Piękos, Piotr, et al.
Publicado: (2025)
Cognitive Load-Aware Inference: A Neuro-Symbolic Framework for Optimizing the Token Economy of Large Language Models
por: Zhang, Yilun
Publicado: (2025)
por: Zhang, Yilun
Publicado: (2025)
Support Vector Boosting Machine (SVBM): Enhancing Classification Performance with AdaBoost and Residual Connections
por: Lian, Junbo Jacob
Publicado: (2024)
por: Lian, Junbo Jacob
Publicado: (2024)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
por: Xin, Meng, et al.
Publicado: (2026)
por: Xin, Meng, et al.
Publicado: (2026)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
por: Czakó, Patrik, et al.
Publicado: (2025)
por: Czakó, Patrik, et al.
Publicado: (2025)
CSRA: Controlled Spectral Residual Augmentation for Robust Sepsis Prediction
por: Guo, Honglin, et al.
Publicado: (2026)
por: Guo, Honglin, et al.
Publicado: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
por: Cho, Yoonjun, et al.
Publicado: (2026)
por: Cho, Yoonjun, et al.
Publicado: (2026)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
por: Zhao, Zhixiong, et al.
Publicado: (2026)
por: Zhao, Zhixiong, et al.
Publicado: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
por: Zhang, Shenao, et al.
Publicado: (2024)
por: Zhang, Shenao, et al.
Publicado: (2024)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
por: Bao, Chengzhu, et al.
Publicado: (2026)
por: Bao, Chengzhu, et al.
Publicado: (2026)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026)
por: Zhao, Minda, et al.
Publicado: (2026)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Residual Cross-Attention Transformer-Based Multi-User CSI Feedback with Deep Joint Source-Channel Coding
por: Zhang, Hengwei, et al.
Publicado: (2025)
por: Zhang, Hengwei, et al.
Publicado: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
por: Chen, Mengzhao, et al.
Publicado: (2025)
por: Chen, Mengzhao, et al.
Publicado: (2025)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
por: Li, Junlin, et al.
Publicado: (2026)
por: Li, Junlin, et al.
Publicado: (2026)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
por: Zhang, Jinghe, et al.
Publicado: (2026)
por: Zhang, Jinghe, et al.
Publicado: (2026)
Achieving binary weight and activation for LLMs using Post-Training Quantization
por: Song, Siqing, et al.
Publicado: (2025)
por: Song, Siqing, et al.
Publicado: (2025)
Automated Formalization via Conceptual Retrieval-Augmented LLMs
por: Lu, Wangyue, et al.
Publicado: (2025)
por: Lu, Wangyue, et al.
Publicado: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
por: Ye, Rongguang, et al.
Publicado: (2025)
por: Ye, Rongguang, et al.
Publicado: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
por: Xu, Yinggan, et al.
Publicado: (2026)
por: Xu, Yinggan, et al.
Publicado: (2026)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
por: Su, Le, et al.
Publicado: (2026)
por: Su, Le, et al.
Publicado: (2026)
Ejemplares similares
-
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
por: Liu, Wenyuan, et al.
Publicado: (2025) -
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
por: Luo, Yilun, et al.
Publicado: (2025) -
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
por: Liu, Wenyuan, et al.
Publicado: (2024) -
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
por: Ma, Xindian, et al.
Publicado: (2026) -
FAAR: Format-Aware Adaptive Rounding for NVFP4
por: Li, Hanglin, et al.
Publicado: (2026)