VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Bhope, Rahul Atul, Jayaram, K. R., Muthusamy, Vinod, Kumar, Ritesh, Isahagian, Vatche, Venkatasubramanian, Nalini |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
Trajectory-Informed Memory Generation for Self-Improving Agent Systems
by: Fang, Gaodan, et al.
Published: (2026)
by: Fang, Gaodan, et al.
Published: (2026)
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows
by: Duesterwald, Evelyn, et al.
Published: (2025)
by: Duesterwald, Evelyn, et al.
Published: (2025)
DIM-SUM: Dynamic IMputation for Smart Utility Management
by: Hildebrant, Ryan, et al.
Published: (2025)
by: Hildebrant, Ryan, et al.
Published: (2025)
On Automating Security Policies with Contemporary LLMs
by: Saura, Pablo Fernández, et al.
Published: (2025)
by: Saura, Pablo Fernández, et al.
Published: (2025)
VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language
by: Wan, Zishuo, et al.
Published: (2025)
by: Wan, Zishuo, et al.
Published: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
by: Raja, Rahul, et al.
Published: (2025)
by: Raja, Rahul, et al.
Published: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
by: Xu, Quanxing, et al.
Published: (2026)
by: Xu, Quanxing, et al.
Published: (2026)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
by: Lan, Jian, et al.
Published: (2025)
by: Lan, Jian, et al.
Published: (2025)
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
by: Ben-Ami, Dan, et al.
Published: (2026)
by: Ben-Ami, Dan, et al.
Published: (2026)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
by: Bai, Xiangyu, et al.
Published: (2026)
by: Bai, Xiangyu, et al.
Published: (2026)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
Multimodal Privacy-Preserving Entity Resolution with Fully Homomorphic Encryption
by: Roy, Susim, et al.
Published: (2026)
by: Roy, Susim, et al.
Published: (2026)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
Automated Road Crack Localization for Spatially Guided Highway Maintenance
by: Knoblauch, Steffen, et al.
Published: (2026)
by: Knoblauch, Steffen, et al.
Published: (2026)
Free Form Medical Visual Question Answering in Radiology
by: Narayanan, Abhishek, et al.
Published: (2024)
by: Narayanan, Abhishek, et al.
Published: (2024)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
by: Lim, Qi Zhi, et al.
Published: (2025)
by: Lim, Qi Zhi, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Multimodal Integration of Human-Like Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge
by: Lando, Giuseppe, et al.
Published: (2026)
by: Lando, Giuseppe, et al.
Published: (2026)
Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
by: Li, Peize, et al.
Published: (2024)
by: Li, Peize, et al.
Published: (2024)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
by: Chen, Jin, et al.
Published: (2024)
by: Chen, Jin, et al.
Published: (2024)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
by: Islam, Md Mohaiminul, et al.
Published: (2025)
by: Islam, Md Mohaiminul, et al.
Published: (2025)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
by: Hao, Dongze, et al.
Published: (2024)
by: Hao, Dongze, et al.
Published: (2024)
CHASE: Competing Hypotheses for Ambiguity-Aware Selective Prediction
by: Jhawar, Kartik, et al.
Published: (2026)
by: Jhawar, Kartik, et al.
Published: (2026)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
Similar Items
-
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
by: Bhope, Rahul Atul, et al.
Published: (2025) -
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
by: Bhope, Rahul Atul, et al.
Published: (2025) -
Trajectory-Informed Memory Generation for Self-Improving Agent Systems
by: Fang, Gaodan, et al.
Published: (2026) -
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows
by: Duesterwald, Evelyn, et al.
Published: (2025) -
DIM-SUM: Dynamic IMputation for Smart Utility Management
by: Hildebrant, Ryan, et al.
Published: (2025)