SPEED-Q: Staged Processing with Enhanced Distillation towards Efficient Low-bit On-device VLM Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Tianyu, Zhao, Shanwei, Zhu, Shiai, Ma, Chenguang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning
by: Wang, Haoxuan, et al.
Published: (2024)
by: Wang, Haoxuan, et al.
Published: (2024)
Q-VLM: Post-training Quantization for Large Vision-Language Models
by: Wang, Changyuan, et al.
Published: (2024)
by: Wang, Changyuan, et al.
Published: (2024)
$γ$-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
by: Fatima, Mishal, et al.
Published: (2025)
by: Fatima, Mishal, et al.
Published: (2025)
A Simple Low-bit Quantization Framework for Video Snapshot Compressive Imaging
by: Cao, Miao, et al.
Published: (2024)
by: Cao, Miao, et al.
Published: (2024)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
by: Sharify, Sayeh, et al.
Published: (2026)
by: Sharify, Sayeh, et al.
Published: (2026)
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
by: Li, Ouxiang, et al.
Published: (2025)
by: Li, Ouxiang, et al.
Published: (2025)
FraQAT: Quantization Aware Training with Fractional bits
by: Morreale, Luca, et al.
Published: (2025)
by: Morreale, Luca, et al.
Published: (2025)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers
by: Huang, Feice, et al.
Published: (2025)
by: Huang, Feice, et al.
Published: (2025)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
GenQ: Quantization in Low Data Regimes with Generative Synthetic Data
by: Li, Yuhang, et al.
Published: (2023)
by: Li, Yuhang, et al.
Published: (2023)
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
by: Tang, Siao, et al.
Published: (2026)
by: Tang, Siao, et al.
Published: (2026)
S$^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
by: Chen, Yujie, et al.
Published: (2025)
by: Chen, Yujie, et al.
Published: (2025)
RTF-Q: Efficient Unsupervised Domain Adaptation with Retraining-free Quantization
by: Du, Nanyang, et al.
Published: (2024)
by: Du, Nanyang, et al.
Published: (2024)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
by: Gao, Chang, et al.
Published: (2024)
by: Gao, Chang, et al.
Published: (2024)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding
by: Tatematsu, Fumiya, et al.
Published: (2026)
by: Tatematsu, Fumiya, et al.
Published: (2026)
Detail Consistent Stage-Wise Distillation for Efficient 3D MRI Segmentation
by: Fan, Mengchen, et al.
Published: (2026)
by: Fan, Mengchen, et al.
Published: (2026)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
by: Ning, Zhenyu, et al.
Published: (2025)
by: Ning, Zhenyu, et al.
Published: (2025)
LEAF: Latent Diffusion with Efficient Encoder Distillation for Aligned Features in Medical Image Segmentation
by: Huang, Qilin, et al.
Published: (2025)
by: Huang, Qilin, et al.
Published: (2025)
Data-Augmented Quantization-Aware Knowledge Distillation
by: Kur, Justin, et al.
Published: (2025)
by: Kur, Justin, et al.
Published: (2025)
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training
by: Zhao, Xiaochen, et al.
Published: (2025)
by: Zhao, Xiaochen, et al.
Published: (2025)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Q-SNNs: Quantized Spiking Neural Networks
by: Wei, Wenjie, et al.
Published: (2024)
by: Wei, Wenjie, et al.
Published: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
MSDNet: Efficient 4D Radar Super-Resolution via Multi-Stage Distillation
by: Huang, Minqing, et al.
Published: (2025)
by: Huang, Minqing, et al.
Published: (2025)
Decoder-Free Distillation for Quantized Image Restoration
by: Sharif, S. M. A., et al.
Published: (2026)
by: Sharif, S. M. A., et al.
Published: (2026)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
by: He, Yefei, et al.
Published: (2023)
by: He, Yefei, et al.
Published: (2023)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Similar Items
-
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025) -
QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning
by: Wang, Haoxuan, et al.
Published: (2024) -
Q-VLM: Post-training Quantization for Large Vision-Language Models
by: Wang, Changyuan, et al.
Published: (2024) -
$γ$-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
by: Fatima, Mishal, et al.
Published: (2025) -
A Simple Low-bit Quantization Framework for Video Snapshot Compressive Imaging
by: Cao, Miao, et al.
Published: (2024)