Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Haokun, Xu, Haobo, Wu, Yichen, Guo, Ziyu, Zhang, Renrui, Lu, Zhichao, Wei, Ying, Zhang, Qingfu, Sun, Zhenan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
di: Lin, Haokun, et al.
Pubblicazione: (2026)
di: Lin, Haokun, et al.
Pubblicazione: (2026)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
di: Yang, Lianwei, et al.
Pubblicazione: (2025)
di: Yang, Lianwei, et al.
Pubblicazione: (2025)
dVoting: Fast Voting for dLLMs
di: Feng, Sicheng, et al.
Pubblicazione: (2026)
di: Feng, Sicheng, et al.
Pubblicazione: (2026)
DMax: Aggressive Parallel Decoding for dLLMs
di: Chen, Zigeng, et al.
Pubblicazione: (2026)
di: Chen, Zigeng, et al.
Pubblicazione: (2026)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
di: Lin, Haokun, et al.
Pubblicazione: (2025)
di: Lin, Haokun, et al.
Pubblicazione: (2025)
dParallel: Learnable Parallel Decoding for dLLMs
di: Chen, Zigeng, et al.
Pubblicazione: (2025)
di: Chen, Zigeng, et al.
Pubblicazione: (2025)
dMoE: dLLMs with Learnable Block Experts
di: Feng, Sicheng, et al.
Pubblicazione: (2026)
di: Feng, Sicheng, et al.
Pubblicazione: (2026)
Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMs
di: Li, Qi, et al.
Pubblicazione: (2025)
di: Li, Qi, et al.
Pubblicazione: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
STDec: Spatio-Temporal Stability Guided Decoding for dLLMs
di: Chen, Yuzhe, et al.
Pubblicazione: (2026)
di: Chen, Yuzhe, et al.
Pubblicazione: (2026)
EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients
di: Hsieh, He-Yen, et al.
Pubblicazione: (2025)
di: Hsieh, He-Yen, et al.
Pubblicazione: (2025)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
Towards Accurate Post-training Quantization for Diffusion Models
di: Wang, Changyuan, et al.
Pubblicazione: (2023)
di: Wang, Changyuan, et al.
Pubblicazione: (2023)
Towards Accurate Post-training Quantization for Reparameterized Models
di: Zhang, Luoming, et al.
Pubblicazione: (2024)
di: Zhang, Luoming, et al.
Pubblicazione: (2024)
QVD: Post-training Quantization for Video Diffusion Models
di: Tian, Shilong, et al.
Pubblicazione: (2024)
di: Tian, Shilong, et al.
Pubblicazione: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
PQD: Post-training Quantization for Efficient Diffusion Models
di: Ye, Jiaojiao, et al.
Pubblicazione: (2024)
di: Ye, Jiaojiao, et al.
Pubblicazione: (2024)
A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
di: Song, Qingyu, et al.
Pubblicazione: (2025)
di: Song, Qingyu, et al.
Pubblicazione: (2025)
SliderQuant: Accurate Post-Training Quantization for LLMs
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
Achieving binary weight and activation for LLMs using Post-Training Quantization
di: Song, Siqing, et al.
Pubblicazione: (2025)
di: Song, Siqing, et al.
Pubblicazione: (2025)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
di: Wu, Junyi, et al.
Pubblicazione: (2024)
di: Wu, Junyi, et al.
Pubblicazione: (2024)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
di: Lv, Keyu, et al.
Pubblicazione: (2026)
di: Lv, Keyu, et al.
Pubblicazione: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
di: Zhang, Jingxuan, et al.
Pubblicazione: (2026)
di: Zhang, Jingxuan, et al.
Pubblicazione: (2026)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
di: Zhang, Tianao, et al.
Pubblicazione: (2025)
di: Zhang, Tianao, et al.
Pubblicazione: (2025)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
di: Zheng, Zekang, et al.
Pubblicazione: (2025)
di: Zheng, Zekang, et al.
Pubblicazione: (2025)
Integer Scale: A Free Lunch for Faster Fine-grained Quantization of LLMs
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
di: Li, Qingyuan, et al.
Pubblicazione: (2024)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
Selective Focus: Investigating Semantics Sensitivity in Post-training Quantization for Lane Detection
di: Fan, Yunqian, et al.
Pubblicazione: (2024)
di: Fan, Yunqian, et al.
Pubblicazione: (2024)
LightningRL: Breaking the Accuracy-Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning
di: Hu, Yanzhe, et al.
Pubblicazione: (2026)
di: Hu, Yanzhe, et al.
Pubblicazione: (2026)
Pyramid Vector Quantization for LLMs
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
di: Luo, Yingsong, et al.
Pubblicazione: (2024)
di: Luo, Yingsong, et al.
Pubblicazione: (2024)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
BiDM: Pushing the Limit of Quantization for Diffusion Models
di: Zheng, Xingyu, et al.
Pubblicazione: (2024)
di: Zheng, Xingyu, et al.
Pubblicazione: (2024)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
di: Lin, Haokun, et al.
Pubblicazione: (2026) -
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
di: Lin, Haokun, et al.
Pubblicazione: (2024) -
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
di: Yang, Lianwei, et al.
Pubblicazione: (2024) -
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
di: Yang, Lianwei, et al.
Pubblicazione: (2025) -
dVoting: Fast Voting for dLLMs
di: Feng, Sicheng, et al.
Pubblicazione: (2026)