Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kurz, Paul Jonas, Wieczorek, Tobias Jan, Abdelsalam, Mohamed A., Aljundi, Rahaf, Rohrbach, Marcus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
Overcoming Generic Knowledge Loss with Selective Parameter Update
by: Zhang, Wenxuan, et al.
Published: (2023)
by: Zhang, Wenxuan, et al.
Published: (2023)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
Efficient Few-Shot Continual Learning in Vision-Language Models
by: Panos, Aristeidis, et al.
Published: (2025)
by: Panos, Aristeidis, et al.
Published: (2025)
Incremental Object-Based Novelty Detection with Feedback Loop
by: Caldarella, Simone, et al.
Published: (2023)
by: Caldarella, Simone, et al.
Published: (2023)
Ego: Embedding-Guided Personalization of Vision-Language Models
by: Seifi, Soroush, et al.
Published: (2026)
by: Seifi, Soroush, et al.
Published: (2026)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
by: Dorovatas, Vaggelis, et al.
Published: (2025)
by: Dorovatas, Vaggelis, et al.
Published: (2025)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026)
by: Rodriguez, Hector G., et al.
Published: (2026)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
Online In-Context Distillation for Low-Resource Vision Language Models
by: Kang, Zhiqi, et al.
Published: (2025)
by: Kang, Zhiqi, et al.
Published: (2025)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
by: Caldarella, Simone, et al.
Published: (2024)
by: Caldarella, Simone, et al.
Published: (2024)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
by: Panos, Aristeidis, et al.
Published: (2024)
by: Panos, Aristeidis, et al.
Published: (2024)
Post-Training Quantization for Video Matting
by: Zhu, Tianrui, et al.
Published: (2025)
by: Zhu, Tianrui, et al.
Published: (2025)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
by: Kou, Qian, et al.
Published: (2026)
by: Kou, Qian, et al.
Published: (2026)
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
by: Seifi, Soroush, et al.
Published: (2025)
by: Seifi, Soroush, et al.
Published: (2025)
Annotation Free Semantic Segmentation with Vision Foundation Models
by: Seifi, Soroush, et al.
Published: (2024)
by: Seifi, Soroush, et al.
Published: (2024)
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
by: Bazi, Yakoub, et al.
Published: (2026)
by: Bazi, Yakoub, et al.
Published: (2026)
Post Training Quantization for Efficient Dataset Condensation
by: Tran, Linh-Tam, et al.
Published: (2026)
by: Tran, Linh-Tam, et al.
Published: (2026)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025)
by: Zohrabi, Reihaneh, et al.
Published: (2025)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
by: Xiang, Ziwei, et al.
Published: (2026)
by: Xiang, Ziwei, et al.
Published: (2026)
QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
Chrono: A Simple Blueprint for Representing Time in MLLMs
by: Rodriguez, Hector, et al.
Published: (2024)
by: Rodriguez, Hector, et al.
Published: (2024)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
First Session Adaptation: A Strong Replay-Free Baseline for Class-Incremental Learning
by: Panos, Aristeidis, et al.
Published: (2023)
by: Panos, Aristeidis, et al.
Published: (2023)
Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification
by: Wu, Xixian, et al.
Published: (2025)
by: Wu, Xixian, et al.
Published: (2025)
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
by: Liu, Xuewen, et al.
Published: (2026)
by: Liu, Xuewen, et al.
Published: (2026)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
by: Bouguerra, Aymen, et al.
Published: (2025)
by: Bouguerra, Aymen, et al.
Published: (2025)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
by: Joshi, Soham, et al.
Published: (2025)
by: Joshi, Soham, et al.
Published: (2025)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
by: Ding, Rui, et al.
Published: (2024)
by: Ding, Rui, et al.
Published: (2024)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)
by: Kim, Gihwan, et al.
Published: (2025)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
by: Chen, Zhuokun, et al.
Published: (2024)
by: Chen, Zhuokun, et al.
Published: (2024)
Similar Items
-
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025) -
Overcoming Generic Knowledge Loss with Selective Parameter Update
by: Zhang, Wenxuan, et al.
Published: (2023) -
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026) -
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)