Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kurz, Paul Jonas, Wieczorek, Tobias Jan, Abdelsalam, Mohamed A., Aljundi, Rahaf, Rohrbach, Marcus
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915797834661888
author Kurz, Paul Jonas
Wieczorek, Tobias Jan
Abdelsalam, Mohamed A.
Aljundi, Rahaf
Rohrbach, Marcus
author_facet Kurz, Paul Jonas
Wieczorek, Tobias Jan
Abdelsalam, Mohamed A.
Aljundi, Rahaf
Rohrbach, Marcus
contents Multimodal Large Language Models (MLLM) are increasingly deployed in domains where both reliability and efficiency are critical. However, current models remain overconfident, producing highly certain but incorrect answers. At the same time, their large size limits deployment on edge devices, necessitating compression. We study the intersection of these two challenges by analyzing how Post-Training Quantization (PTQ) compression affects both accuracy and reliability in Visual Question Answering (VQA). We evaluate two MLLMs, Qwen2-VL-7B and Idefics3-8B, quantized with data-free (HQQ) and data-aware (MBQ) methods across multiple bit widths. To counteract the reduction in reliability caused by quantization, we adapt the Selector confidence estimator for quantized multimodal settings and test its robustness across various quantization levels and out-of-distribution (OOD) scenarios. We find that PTQ degrades both accuracy and reliability. Data-aware methods soften the effect thereof. The Selector substantially mitigates the reliability impact. The combination of int4 MBQ and the Selector achieves the best efficiency-reliability trade-off, closing in on uncompressed performance at approx. 75% less memory demand. Overall, we present the first systematic study linking quantization and reliability in multimodal settings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13289
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
Kurz, Paul Jonas
Wieczorek, Tobias Jan
Abdelsalam, Mohamed A.
Aljundi, Rahaf
Rohrbach, Marcus
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal Large Language Models (MLLM) are increasingly deployed in domains where both reliability and efficiency are critical. However, current models remain overconfident, producing highly certain but incorrect answers. At the same time, their large size limits deployment on edge devices, necessitating compression. We study the intersection of these two challenges by analyzing how Post-Training Quantization (PTQ) compression affects both accuracy and reliability in Visual Question Answering (VQA). We evaluate two MLLMs, Qwen2-VL-7B and Idefics3-8B, quantized with data-free (HQQ) and data-aware (MBQ) methods across multiple bit widths. To counteract the reduction in reliability caused by quantization, we adapt the Selector confidence estimator for quantized multimodal settings and test its robustness across various quantization levels and out-of-distribution (OOD) scenarios. We find that PTQ degrades both accuracy and reliability. Data-aware methods soften the effect thereof. The Selector substantially mitigates the reliability impact. The combination of int4 MBQ and the Selector achieves the best efficiency-reliability trade-off, closing in on uncompressed performance at approx. 75% less memory demand. Overall, we present the first systematic study linking quantization and reliability in multimodal settings.
title Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.13289