Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Xijie, Shen, Zhiqiang, Dong, Pingcheng, Cheng, Kwang-Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
von: Su, Le, et al.
Veröffentlicht: (2026)
von: Su, Le, et al.
Veröffentlicht: (2026)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
von: Ding, Rui, et al.
Veröffentlicht: (2024)
von: Ding, Rui, et al.
Veröffentlicht: (2024)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
QuEPT: Quantized Elastic Precision Transformers with One-Shot Calibration for Multi-Bit Switching
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective
von: Yin, Zeyuan, et al.
Veröffentlicht: (2023)
von: Yin, Zeyuan, et al.
Veröffentlicht: (2023)
Memory Efficient Transformer Adapter for Dense Predictions
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
When Bits Break Recourse: Counterfactual-Faithful Quantization
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization
von: Luo, Róisín, et al.
Veröffentlicht: (2024)
von: Luo, Róisín, et al.
Veröffentlicht: (2024)
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
von: Bhatnagar, Shubhang, et al.
Veröffentlicht: (2025)
von: Bhatnagar, Shubhang, et al.
Veröffentlicht: (2025)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
DELT: A Simple Diversity-driven EarlyLate Training for Dataset Distillation
von: Shen, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Shen, Zhiqiang, et al.
Veröffentlicht: (2024)
Amortized-Precision Quantization for Early-Exit Vision Transformers
von: Fang, Rui, et al.
Veröffentlicht: (2026)
von: Fang, Rui, et al.
Veröffentlicht: (2026)
From Fewer Samples to Fewer Bits: Reframing Dataset Distillation as Joint Optimization of Precision and Compactness
von: Dinh, My H., et al.
Veröffentlicht: (2026)
von: Dinh, My H., et al.
Veröffentlicht: (2026)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
von: Matişan, Răzvan-Andrei, et al.
Veröffentlicht: (2025)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
von: Spadaro, Gabriele, et al.
Veröffentlicht: (2025)
von: Spadaro, Gabriele, et al.
Veröffentlicht: (2025)
CMOSE: Comprehensive Multi-Modality Online Student Engagement Dataset with High-Quality Labels
von: Wu, Chi-hsuan, et al.
Veröffentlicht: (2023)
von: Wu, Chi-hsuan, et al.
Veröffentlicht: (2023)
Dataset Distillation via Curriculum Data Synthesis in Large Data Era
von: Yin, Zeyuan, et al.
Veröffentlicht: (2023)
von: Yin, Zeyuan, et al.
Veröffentlicht: (2023)
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
von: Zhang, Kevin, et al.
Veröffentlicht: (2022)
von: Zhang, Kevin, et al.
Veröffentlicht: (2022)
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
Towards Precise Scaling Laws for Video Diffusion Transformers
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
When Training-Free NAS Meets Vision Transformer: A Neural Tangent Kernel Perspective
von: Zhou, Qiqi, et al.
Veröffentlicht: (2024)
von: Zhou, Qiqi, et al.
Veröffentlicht: (2024)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
von: Lee, Dongyeun, et al.
Veröffentlicht: (2025)
von: Lee, Dongyeun, et al.
Veröffentlicht: (2025)
HVI: A New Color Space for Low-light Image Enhancement
von: Yan, Qingsen, et al.
Veröffentlicht: (2025)
von: Yan, Qingsen, et al.
Veröffentlicht: (2025)
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
von: Du, Jinyang, et al.
Veröffentlicht: (2026)
von: Du, Jinyang, et al.
Veröffentlicht: (2026)
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
von: Zhang, Jinhao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinhao, et al.
Veröffentlicht: (2025)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
Extreme Model Compression with Structured Sparsity at Low Precision
von: Liu, Dan, et al.
Veröffentlicht: (2025)
von: Liu, Dan, et al.
Veröffentlicht: (2025)
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
von: Bouguerra, Aymen, et al.
Veröffentlicht: (2025)
von: Bouguerra, Aymen, et al.
Veröffentlicht: (2025)
Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models
von: Zou, Jade, et al.
Veröffentlicht: (2026)
von: Zou, Jade, et al.
Veröffentlicht: (2026)
OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
von: Chu, Huanpeng, et al.
Veröffentlicht: (2025)
von: Chu, Huanpeng, et al.
Veröffentlicht: (2025)
MixMask: Revisiting Masking Strategy for Siamese ConvNets
von: Vishniakov, Kirill, et al.
Veröffentlicht: (2022)
von: Vishniakov, Kirill, et al.
Veröffentlicht: (2022)
Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
von: Liu, Shih-yang, et al.
Veröffentlicht: (2023) -
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
von: Huang, Xijie, et al.
Veröffentlicht: (2023) -
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
von: Su, Le, et al.
Veröffentlicht: (2026) -
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
von: Ding, Rui, et al.
Veröffentlicht: (2024) -
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
von: Choi, Kanghyun, et al.
Veröffentlicht: (2024)