Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Euntae, Song, Sumin, Lim, Woosang, Yoo, Sungjoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025)
by: Son, Donghyun, et al.
Published: (2025)
LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss
by: Choi, Euntae, et al.
Published: (2026)
by: Choi, Euntae, et al.
Published: (2026)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
by: Czakó, Patrik, et al.
Published: (2025)
by: Czakó, Patrik, et al.
Published: (2025)
BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
by: He, Liulu, et al.
Published: (2025)
by: He, Liulu, et al.
Published: (2025)
Group Sequence Policy Optimization
by: Zheng, Chujie, et al.
Published: (2025)
by: Zheng, Chujie, et al.
Published: (2025)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Efficient and Scalable Estimation of Tool Representations in Vector Space
by: Moon, Suhong, et al.
Published: (2024)
by: Moon, Suhong, et al.
Published: (2024)
ARGUS: Adaptive Rotation-Invariant Geometric Unsupervised System
by: Sharma, Anantha
Published: (2026)
by: Sharma, Anantha
Published: (2026)
Parameter Efficient Quasi-Orthogonal Fine-Tuning via Givens Rotation
by: Ma, Xinyu, et al.
Published: (2024)
by: Ma, Xinyu, et al.
Published: (2024)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
by: Lim, Woosang, et al.
Published: (2025)
by: Lim, Woosang, et al.
Published: (2025)
Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
by: Yoo, YoungJoon, et al.
Published: (2023)
by: Yoo, YoungJoon, et al.
Published: (2023)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
by: Jiang, Yichen, et al.
Published: (2024)
by: Jiang, Yichen, et al.
Published: (2024)
Encoding Agent Trajectories as Representations with Sequence Transformers
by: Tsiligkaridis, Athanasios, et al.
Published: (2024)
by: Tsiligkaridis, Athanasios, et al.
Published: (2024)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
by: Cheong, Minsoo, et al.
Published: (2026)
by: Cheong, Minsoo, et al.
Published: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
Phys3DGS: Physically-based 3D Gaussian Splatting for Inverse Rendering
by: Choi, Euntae, et al.
Published: (2024)
by: Choi, Euntae, et al.
Published: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
by: Xu, Bingxin, et al.
Published: (2025)
by: Xu, Bingxin, et al.
Published: (2025)
C$^2$GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
by: Liu, Haotian, et al.
Published: (2025)
by: Liu, Haotian, et al.
Published: (2025)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
LLM4Delay: Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
by: Phisannupawong, Thaweerath, et al.
Published: (2025)
by: Phisannupawong, Thaweerath, et al.
Published: (2025)
Optimizing Large Language Models through Quantization: A Comparative Analysis of PTQ and QAT Techniques
by: Hasan, Jahid
Published: (2024)
by: Hasan, Jahid
Published: (2024)
DNAZEN: Enhanced Gene Sequence Representations via Mixed Granularities of Coding Units
by: Mao, Lei, et al.
Published: (2025)
by: Mao, Lei, et al.
Published: (2025)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
Strategic Fusion Optimizes Transformer Compression
by: Rahman, Md Shoaibur
Published: (2025)
by: Rahman, Md Shoaibur
Published: (2025)
Momentum Streams for Optimizer-Inspired Transformers
by: Gai, Jingchu, et al.
Published: (2026)
by: Gai, Jingchu, et al.
Published: (2026)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
by: Song, Yurun, et al.
Published: (2025)
by: Song, Yurun, et al.
Published: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
by: Wang, Jialu, et al.
Published: (2026)
by: Wang, Jialu, et al.
Published: (2026)
On the Geometric Structure of Layer Updates in Deep Language Models
by: Yoo, Jun-Sik
Published: (2026)
by: Yoo, Jun-Sik
Published: (2026)
Similar Items
-
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025) -
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025) -
LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss
by: Choi, Euntae, et al.
Published: (2026) -
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025) -
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
by: Su, Zunhai, et al.
Published: (2025)