Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhaoyang, Wang, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization
von: Luo, Róisín, et al.
Veröffentlicht: (2024)
von: Luo, Róisín, et al.
Veröffentlicht: (2024)
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
von: Ding, Rui, et al.
Veröffentlicht: (2024)
von: Ding, Rui, et al.
Veröffentlicht: (2024)
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
von: Du, Jinyang, et al.
Veröffentlicht: (2026)
von: Du, Jinyang, et al.
Veröffentlicht: (2026)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
von: Su, Le, et al.
Veröffentlicht: (2026)
von: Su, Le, et al.
Veröffentlicht: (2026)
Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
von: Frumkin, Natalia, et al.
Veröffentlicht: (2025)
von: Frumkin, Natalia, et al.
Veröffentlicht: (2025)
QuEPT: Quantized Elastic Precision Transformers with One-Shot Calibration for Multi-Bit Switching
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
Q-SAM2: Accurate Quantization for Segment Anything Model 2
von: Farronato, Nicola, et al.
Veröffentlicht: (2025)
von: Farronato, Nicola, et al.
Veröffentlicht: (2025)
P4Q: Learning to Prompt for Quantization in Visual-language Models
von: Sun, Huixin, et al.
Veröffentlicht: (2024)
von: Sun, Huixin, et al.
Veröffentlicht: (2024)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
When Bits Break Recourse: Counterfactual-Faithful Quantization
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
Channel-wise Vector Quantization
von: Song, Wei, et al.
Veröffentlicht: (2026)
von: Song, Wei, et al.
Veröffentlicht: (2026)
Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning
von: Ma, Lianbo, et al.
Veröffentlicht: (2025)
von: Ma, Lianbo, et al.
Veröffentlicht: (2025)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
von: Bhatnagar, Shubhang, et al.
Veröffentlicht: (2025)
von: Bhatnagar, Shubhang, et al.
Veröffentlicht: (2025)
Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
Scalable Image Tokenization with Index Backpropagation Quantization
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
Scaling Image Tokenizers with Grouped Spherical Quantization
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
von: Zhang, Jiaji, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaji, et al.
Veröffentlicht: (2025)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Self-Supervised Quantization-Aware Knowledge Distillation
von: Zhao, Kaiqi, et al.
Veröffentlicht: (2024)
von: Zhao, Kaiqi, et al.
Veröffentlicht: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
von: Wu, Junhao, et al.
Veröffentlicht: (2026)
von: Wu, Junhao, et al.
Veröffentlicht: (2026)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
von: Lee, Jemin, et al.
Veröffentlicht: (2023)
von: Lee, Jemin, et al.
Veröffentlicht: (2023)
Efficient Quantization-Aware Training on Segment Anything Model in Medical Images and Its Deployment
von: Lu, Haisheng, et al.
Veröffentlicht: (2024)
von: Lu, Haisheng, et al.
Veröffentlicht: (2024)
AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation
von: Deb, Prantik, et al.
Veröffentlicht: (2026)
von: Deb, Prantik, et al.
Veröffentlicht: (2026)
Quantization-Aware Neuromorphic Architecture for Skin Disease Classification on Resource-Constrained Devices
von: Wang, Haitian, et al.
Veröffentlicht: (2025)
von: Wang, Haitian, et al.
Veröffentlicht: (2025)
PTQAT: A Hybrid Parameter-Efficient Quantization Algorithm for 3D Perception Tasks
von: Wang, Xinhao, et al.
Veröffentlicht: (2025)
von: Wang, Xinhao, et al.
Veröffentlicht: (2025)
Post-Training Quantization for Video Matting
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
von: Maisonnave, Lucas, et al.
Veröffentlicht: (2025)
DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization
von: Wu, Qiuxia, et al.
Veröffentlicht: (2025)
von: Wu, Qiuxia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024) -
Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization
von: Luo, Róisín, et al.
Veröffentlicht: (2024) -
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026) -
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
von: Zhong, Yi, et al.
Veröffentlicht: (2026) -
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
von: Huang, Xijie, et al.
Veröffentlicht: (2023)