Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Zhanfeng, Guo, Shuai, Di, Xin, Peng, Long, Cao, Yang, Zha, Zhengjun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiFloat4 Format for Language Model Inference
by: Luo, Yuanyong, et al.
Published: (2026)
by: Luo, Yuanyong, et al.
Published: (2026)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026)
by: Taghian, Mehran, et al.
Published: (2026)
Ascend HiFloat8 Format for Deep Learning
by: Luo, Yuanyong, et al.
Published: (2024)
by: Luo, Yuanyong, et al.
Published: (2024)
Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
by: Wu, Junhao, et al.
Published: (2026)
by: Wu, Junhao, et al.
Published: (2026)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
by: Zhao, Yiming
Published: (2026)
by: Zhao, Yiming
Published: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting
by: Shi, Mingyu, et al.
Published: (2026)
by: Shi, Mingyu, et al.
Published: (2026)
Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image Restoration
by: Peng, Long, et al.
Published: (2025)
by: Peng, Long, et al.
Published: (2025)
Tail-Aware Post-Training Quantization for 3D Geometry Models
by: Pan, Sicheng, et al.
Published: (2026)
by: Pan, Sicheng, et al.
Published: (2026)
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
by: Du, Jinyang, et al.
Published: (2026)
by: Du, Jinyang, et al.
Published: (2026)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
PTQ4SAM: Post-Training Quantization for Segment Anything
by: Lv, Chengtao, et al.
Published: (2024)
by: Lv, Chengtao, et al.
Published: (2024)
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers
by: Chen, Ruichen, et al.
Published: (2025)
by: Chen, Ruichen, et al.
Published: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
decoupleQ: Towards 2-bit Post-Training Uniform Quantization via decoupling Parameters into Integer and Floating Points
by: Guo, Yi, et al.
Published: (2024)
by: Guo, Yi, et al.
Published: (2024)
PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement
by: Feng, ZhanFeng, et al.
Published: (2025)
by: Feng, ZhanFeng, et al.
Published: (2025)
PTQ4RIS: Post-Training Quantization for Referring Image Segmentation
by: Jiang, Xiaoyan, et al.
Published: (2024)
by: Jiang, Xiaoyan, et al.
Published: (2024)
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
by: Wang, Chengfeng, et al.
Published: (2025)
by: Wang, Chengfeng, et al.
Published: (2025)
PTQ4VM: Post-Training Quantization for Visual Mamba
by: Cho, Younghyun, et al.
Published: (2024)
by: Cho, Younghyun, et al.
Published: (2024)
EMoTive: Event-guided Trajectory Modeling for 3D Motion Estimation
by: Wan, Zengyu, et al.
Published: (2025)
by: Wan, Zengyu, et al.
Published: (2025)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers
by: Deng, Juncan, et al.
Published: (2024)
by: Deng, Juncan, et al.
Published: (2024)
2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
Max-Linear Tail Regression
by: Chen, Liujun, et al.
Published: (2025)
by: Chen, Liujun, et al.
Published: (2025)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
Efficient Quantization-Aware Neural Receivers: Beyond Post-Training Quantization
by: Yellapragada, SaiKrishna Saketh, et al.
Published: (2025)
by: Yellapragada, SaiKrishna Saketh, et al.
Published: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
by: Yu, Xiaoming, et al.
Published: (2026)
by: Yu, Xiaoming, et al.
Published: (2026)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
by: Blumenberg, Patrick, et al.
Published: (2025)
by: Blumenberg, Patrick, et al.
Published: (2025)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
by: Liu, Xuewen, et al.
Published: (2026)
by: Liu, Xuewen, et al.
Published: (2026)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
P$^2$-ViT: Power-of-Two Post-Training Quantization and Acceleration for Fully Quantized Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
by: Xin, Meng, et al.
Published: (2026)
by: Xin, Meng, et al.
Published: (2026)
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization
by: Zhang, Aozhong, et al.
Published: (2024)
by: Zhang, Aozhong, et al.
Published: (2024)
Weight Group-wise Post-Training Quantization for Medical Foundation Model
by: Chen, Yineng, et al.
Published: (2026)
by: Chen, Yineng, et al.
Published: (2026)
PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
by: Vora, Jayneel, et al.
Published: (2024)
by: Vora, Jayneel, et al.
Published: (2024)
Similar Items
-
HiFloat4 Format for Language Model Inference
by: Luo, Yuanyong, et al.
Published: (2026) -
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026) -
Ascend HiFloat8 Format for Deep Learning
by: Luo, Yuanyong, et al.
Published: (2024) -
Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
by: Wu, Junhao, et al.
Published: (2026) -
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
by: Zhao, Yiming
Published: (2026)