VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhope, Rahul Atul, Jayaram, K. R., Muthusamy, Vinod, Kumar, Ritesh, Isahagian, Vatche, Venkatasubramanian, Nalini
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910009916391424
author Bhope, Rahul Atul
Jayaram, K. R.
Muthusamy, Vinod
Kumar, Ritesh
Isahagian, Vatche
Venkatasubramanian, Nalini
author_facet Bhope, Rahul Atul
Jayaram, K. R.
Muthusamy, Vinod
Kumar, Ritesh
Isahagian, Vatche
Venkatasubramanian, Nalini
contents Despite significant costs from retrieving and processing high-fidelity visual inputs, most multimodal vision-language systems operate at fixed fidelity levels. We introduce VOILA, a framework for Value-Of-Information-driven adaptive fidelity selection in Visual Question Answering (VQA) that optimizes what information to retrieve before model execution. Given a query, VOILA uses a two-stage pipeline: a gradient-boosted regressor estimates correctness likelihood at each fidelity from question features alone, then an isotonic calibrator refines these probabilities for reliable decision-making. The system selects the minimum-cost fidelity maximizing expected utility given predicted accuracy and retrieval costs. We evaluate VOILA across three deployment scenarios using five datasets (VQA-v2, GQA, TextVQA, LoCoMo, FloodNet) and six Vision-Language Models (VLMs) with 7B-235B parameters. VOILA consistently achieves 50-60% cost reductions while retaining 90-95% of full-resolution accuracy across diverse query types and model architectures, demonstrating that pre-retrieval fidelity selection is vital to optimize multimodal inference under resource constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03007
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
Bhope, Rahul Atul
Jayaram, K. R.
Muthusamy, Vinod
Kumar, Ritesh
Isahagian, Vatche
Venkatasubramanian, Nalini
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Despite significant costs from retrieving and processing high-fidelity visual inputs, most multimodal vision-language systems operate at fixed fidelity levels. We introduce VOILA, a framework for Value-Of-Information-driven adaptive fidelity selection in Visual Question Answering (VQA) that optimizes what information to retrieve before model execution. Given a query, VOILA uses a two-stage pipeline: a gradient-boosted regressor estimates correctness likelihood at each fidelity from question features alone, then an isotonic calibrator refines these probabilities for reliable decision-making. The system selects the minimum-cost fidelity maximizing expected utility given predicted accuracy and retrieval costs. We evaluate VOILA across three deployment scenarios using five datasets (VQA-v2, GQA, TextVQA, LoCoMo, FloodNet) and six Vision-Language Models (VLMs) with 7B-235B parameters. VOILA consistently achieves 50-60% cost reductions while retaining 90-95% of full-resolution accuracy across diverse query types and model architectures, demonstrating that pre-retrieval fidelity selection is vital to optimize multimodal inference under resource constraints.
title VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.03007