MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Seoungsub, Kim, In Seo, Kim, Seon Wook |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
por: Cho, Yoonjun, et al.
Publicado: (2025)
por: Cho, Yoonjun, et al.
Publicado: (2025)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
por: Zhang, Stephen, et al.
Publicado: (2024)
por: Zhang, Stephen, et al.
Publicado: (2024)
Towards Scalable Handwriting Communication via EEG Decoding and Latent Embedding Integration
por: Kim, Jun-Young, et al.
Publicado: (2024)
por: Kim, Jun-Young, et al.
Publicado: (2024)
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
por: Kim, Seung-Wook, et al.
Publicado: (2025)
por: Kim, Seung-Wook, et al.
Publicado: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
por: Federici, Marco, et al.
Publicado: (2025)
por: Federici, Marco, et al.
Publicado: (2025)
ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models
por: Kim, Dongha, et al.
Publicado: (2023)
por: Kim, Dongha, et al.
Publicado: (2023)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)
por: Yang, June Yong, et al.
Publicado: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
por: Lee, Dongyeun, et al.
Publicado: (2025)
por: Lee, Dongyeun, et al.
Publicado: (2025)
Compressing Large Language Models using Low Rank and Low Precision Decomposition
por: Saha, Rajarshi, et al.
Publicado: (2024)
por: Saha, Rajarshi, et al.
Publicado: (2024)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
por: Lee, Jung Hyun, et al.
Publicado: (2024)
por: Lee, Jung Hyun, et al.
Publicado: (2024)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
por: Cho, Yoonjun, et al.
Publicado: (2026)
por: Cho, Yoonjun, et al.
Publicado: (2026)
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
por: Zhan, Xiaohua, et al.
Publicado: (2026)
por: Zhan, Xiaohua, et al.
Publicado: (2026)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
por: Rakka, Mariam, et al.
Publicado: (2025)
por: Rakka, Mariam, et al.
Publicado: (2025)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
por: An, Selim, et al.
Publicado: (2026)
por: An, Selim, et al.
Publicado: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
por: Bouzouad, Meriem, et al.
Publicado: (2026)
por: Bouzouad, Meriem, et al.
Publicado: (2026)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
por: Liu, Wenyuan, et al.
Publicado: (2025)
por: Liu, Wenyuan, et al.
Publicado: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
Low-Rank Tensor Decompositions for the Theory of Neural Networks
por: Borsoi, Ricardo, et al.
Publicado: (2025)
por: Borsoi, Ricardo, et al.
Publicado: (2025)
Random Conditioning with Distillation for Data-Efficient Diffusion Model Compression
por: Kim, Dohyun, et al.
Publicado: (2025)
por: Kim, Dohyun, et al.
Publicado: (2025)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
por: Zhang, Tao, et al.
Publicado: (2025)
por: Zhang, Tao, et al.
Publicado: (2025)
Decoupling General and Personalized Knowledge in Federated Learning via Additive and Low-Rank Decomposition
por: Wu, Xinghao, et al.
Publicado: (2024)
por: Wu, Xinghao, et al.
Publicado: (2024)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
por: Ham, Seokil, et al.
Publicado: (2024)
por: Ham, Seokil, et al.
Publicado: (2024)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
por: Zhang, Shu-Hao, et al.
Publicado: (2026)
por: Zhang, Shu-Hao, et al.
Publicado: (2026)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
por: Lee, Geonho, et al.
Publicado: (2024)
por: Lee, Geonho, et al.
Publicado: (2024)
Low-Rank Quantization-Aware Training for LLMs
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
por: Varshney, Ayush K., et al.
Publicado: (2026)
por: Varshney, Ayush K., et al.
Publicado: (2026)
Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs
por: Kaliaperumal, Pranav Kumar
Publicado: (2026)
por: Kaliaperumal, Pranav Kumar
Publicado: (2026)
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
por: Jung, Yeonjoon, et al.
Publicado: (2025)
por: Jung, Yeonjoon, et al.
Publicado: (2025)
Optimal Policy Sparsification and Low Rank Decomposition for Deep Reinforcement Learning
por: Goddla, Vikram
Publicado: (2024)
por: Goddla, Vikram
Publicado: (2024)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
por: Kim, Junhan, et al.
Publicado: (2026)
por: Kim, Junhan, et al.
Publicado: (2026)
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
por: Luo, Haozheng, et al.
Publicado: (2025)
por: Luo, Haozheng, et al.
Publicado: (2025)
BoA: Attention-aware Post-training Quantization without Backpropagation
por: Kim, Junhan, et al.
Publicado: (2024)
por: Kim, Junhan, et al.
Publicado: (2024)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
por: Zhang, Zhiyuan, et al.
Publicado: (2026)
por: Zhang, Zhiyuan, et al.
Publicado: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
por: Huang, Wei, et al.
Publicado: (2023)
por: Huang, Wei, et al.
Publicado: (2023)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
por: Kim, Junhan, et al.
Publicado: (2024)
por: Kim, Junhan, et al.
Publicado: (2024)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
por: Abillama, Pierre, et al.
Publicado: (2025)
por: Abillama, Pierre, et al.
Publicado: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
por: Tang, Pingzhi, et al.
Publicado: (2026)
por: Tang, Pingzhi, et al.
Publicado: (2026)
Ejemplares similares
-
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
por: Cho, Yoonjun, et al.
Publicado: (2025) -
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
por: Zhang, Stephen, et al.
Publicado: (2024) -
Towards Scalable Handwriting Communication via EEG Decoding and Latent Embedding Integration
por: Kim, Jun-Young, et al.
Publicado: (2024) -
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
por: Kim, Seung-Wook, et al.
Publicado: (2025) -
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
por: Federici, Marco, et al.
Publicado: (2025)