VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Yufei, Huang, Yushi, Shao, Jiawei, Zhu, Lunjie, Zhang, Chi, Li, Xuelong, Zhang, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
Towards Accurate Post-Training Quantization of Vision Transformers via Error Reduction
von: Zhong, Yunshan, et al.
Veröffentlicht: (2024)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2024)
Dynamic Token Reweighting for Robust Vision-Language Models
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2025)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2025)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
Enhancing Post-Training Quantization via Future Activation Awareness
von: Lv, Zheqi, et al.
Veröffentlicht: (2026)
von: Lv, Zheqi, et al.
Veröffentlicht: (2026)
Real-Time Human Frontal View Synthesis from a Single Image
von: Lin, Fangyu, et al.
Veröffentlicht: (2026)
von: Lin, Fangyu, et al.
Veröffentlicht: (2026)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
Privacy-Aware Camera 2.0 Technical Report
von: Song, Huan, et al.
Veröffentlicht: (2026)
von: Song, Huan, et al.
Veröffentlicht: (2026)
AdaLog: Post-Training Quantization for Vision Transformers with Adaptive Logarithm Quantizer
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2024)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers
von: Jiang, Runqing, et al.
Veröffentlicht: (2025)
von: Jiang, Runqing, et al.
Veröffentlicht: (2025)
LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
von: Li, Xue, et al.
Veröffentlicht: (2025)
von: Li, Xue, et al.
Veröffentlicht: (2025)
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
von: Zhang, Jipeng, et al.
Veröffentlicht: (2025)
von: Zhang, Jipeng, et al.
Veröffentlicht: (2025)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
von: Meituan LongCat Team, et al.
Veröffentlicht: (2026)
von: Meituan LongCat Team, et al.
Veröffentlicht: (2026)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
von: Cheng, Liwei, et al.
Veröffentlicht: (2026)
von: Cheng, Liwei, et al.
Veröffentlicht: (2026)
APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
von: Liang, Guang, et al.
Veröffentlicht: (2025)
von: Liang, Guang, et al.
Veröffentlicht: (2025)
Vision-centric Token Compression in Large Language Model
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
Training-free Token Reduction for Vision Mamba
von: Ma, Qiankun, et al.
Veröffentlicht: (2025)
von: Ma, Qiankun, et al.
Veröffentlicht: (2025)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
von: Shao, Zhenwei, et al.
Veröffentlicht: (2025)
von: Shao, Zhenwei, et al.
Veröffentlicht: (2025)
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026) -
Towards Accurate Post-Training Quantization of Vision Transformers via Error Reduction
von: Zhong, Yunshan, et al.
Veröffentlicht: (2024) -
Dynamic Token Reweighting for Robust Vision-Language Models
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2025) -
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026) -
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)