ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Thrash, Chayne, Abbasi, Ali, Kolouri, Soheil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025)
by: Qin, Haoran, et al.
Published: (2025)
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
MCNC: Manifold-Constrained Reparameterization for Neural Compression
by: Thrash, Chayne, et al.
Published: (2024)
by: Thrash, Chayne, et al.
Published: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
One Category One Prompt: Dataset Distillation using Diffusion Models
by: Abbasi, Ali, et al.
Published: (2024)
by: Abbasi, Ali, et al.
Published: (2024)
Physics informed cell representations for variational formulation of multiscale problems
by: Gao, Yuxiang, et al.
Published: (2024)
by: Gao, Yuxiang, et al.
Published: (2024)
EMPEROR: Efficient Moment-Preserving Representation of Distributions
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Vector-Quantized Soft Label Compression for Dataset Distillation
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
LUNA: Linear Universal Neural Attention with Generalization Guarantees
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Sinkhorn-Drifting Generative Models
by: He, Ping, et al.
Published: (2026)
by: He, Ping, et al.
Published: (2026)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
by: Shahbazi, Ashkan, et al.
Published: (2026)
by: Shahbazi, Ashkan, et al.
Published: (2026)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
QuIDE: Mastering the Quantized Intelligence Trade-off via Active Optimization
by: Jiang, Xiantao
Published: (2026)
by: Jiang, Xiantao
Published: (2026)
Fused Partial Gromov-Wasserstein for Structured Objects
by: Bai, Yikun, et al.
Published: (2025)
by: Bai, Yikun, et al.
Published: (2025)
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Predicting Probabilities of Error to Combine Quantization and Early Exiting: QuEE
by: Regol, Florence, et al.
Published: (2024)
by: Regol, Florence, et al.
Published: (2024)
Linear Partial Gromov-Wasserstein Embedding
by: Bai, Yikun, et al.
Published: (2024)
by: Bai, Yikun, et al.
Published: (2024)
Constrained Sliced Wasserstein Embedding
by: NaderiAlizadeh, Navid, et al.
Published: (2025)
by: NaderiAlizadeh, Navid, et al.
Published: (2025)
Equivariant vs. Invariant Layers: A Comparison of Backbone and Pooling for Point Cloud Classification
by: Kothapalli, Abihith, et al.
Published: (2023)
by: Kothapalli, Abihith, et al.
Published: (2023)
Partial Gromov-Wasserstein Metric
by: Bai, Yikun, et al.
Published: (2024)
by: Bai, Yikun, et al.
Published: (2024)
DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation
by: Xiang, Jingyang, et al.
Published: (2024)
by: Xiang, Jingyang, et al.
Published: (2024)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
QuATON: Quantization Aware Training of Optical Neurons
by: Kariyawasam, Hasindu, et al.
Published: (2023)
by: Kariyawasam, Hasindu, et al.
Published: (2023)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
Statistical Context Detection for Deep Lifelong Reinforcement Learning
by: Dick, Jeffery, et al.
Published: (2024)
by: Dick, Jeffery, et al.
Published: (2024)
Compander-Aligned Query Geometry for Quantized Zeroth-Order Optimization
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
by: Chee, Jerry, et al.
Published: (2023)
by: Chee, Jerry, et al.
Published: (2023)
QuAILoRA: Quantization-Aware Initialization for LoRA
by: Lawton, Neal, et al.
Published: (2024)
by: Lawton, Neal, et al.
Published: (2024)
Linear Optimal Partial Transport Embedding
by: Bai, Yikun, et al.
Published: (2023)
by: Bai, Yikun, et al.
Published: (2023)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
Optimizing LLMs Using Quantization for Mobile Execution
by: Yadav, Agatsya, et al.
Published: (2025)
by: Yadav, Agatsya, et al.
Published: (2025)
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
by: Zhou, Jiajun, et al.
Published: (2025)
by: Zhou, Jiajun, et al.
Published: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
OT-MeanFlow3D: Bridging Optimal Transport and Meanflow for Efficient 3D Point Cloud Generation
by: Akbari, Elaheh, et al.
Published: (2025)
by: Akbari, Elaheh, et al.
Published: (2025)
QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling
by: Sharma, Preetam, et al.
Published: (2026)
by: Sharma, Preetam, et al.
Published: (2026)
ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models
by: Zou, Xiaocheng, et al.
Published: (2025)
by: Zou, Xiaocheng, et al.
Published: (2025)
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
by: Tran, Huy, et al.
Published: (2024)
by: Tran, Huy, et al.
Published: (2024)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
by: Yang, Jaewoo, et al.
Published: (2024)
by: Yang, Jaewoo, et al.
Published: (2024)
Similar Items
-
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026) -
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025) -
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026) -
MCNC: Manifold-Constrained Reparameterization for Neural Compression
by: Thrash, Chayne, et al.
Published: (2024) -
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
by: Shahbazi, Ashkan, et al.
Published: (2025)