Enhancing Multi-Image Question Answering via Submodular Subset Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Aaryan, Gupta, Shivansh, Agarwal, Samar, C., Vishak Prasad, Ramakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
by: Gupta, Madhav, et al.
Published: (2025)
by: Gupta, Madhav, et al.
Published: (2025)
Less is More: Fewer Interpretable Region via Submodular Subset Selection
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Improving Video Question Answering through query-based frame selection
by: Patil, Himanshu, et al.
Published: (2026)
by: Patil, Himanshu, et al.
Published: (2026)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
by: Chintapatla, Ishant, et al.
Published: (2025)
by: Chintapatla, Ishant, et al.
Published: (2025)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
by: Böther, Maximilian, et al.
Published: (2024)
by: Böther, Maximilian, et al.
Published: (2024)
Describe Anything Model for Visual Question Answering on Text-rich Images
by: Vu, Yen-Linh, et al.
Published: (2025)
by: Vu, Yen-Linh, et al.
Published: (2025)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Lossless Image Compression Using Multi-level Dictionaries: Binary Images
by: Agnihotri, Samar, et al.
Published: (2024)
by: Agnihotri, Samar, et al.
Published: (2024)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
by: Bar, Noga, et al.
Published: (2024)
by: Bar, Noga, et al.
Published: (2024)
Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection
by: Wan, Zhijing, et al.
Published: (2025)
by: Wan, Zhijing, et al.
Published: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
Bandit Guided Submodular Curriculum for Adaptive Subset Selection
by: Chanda, Prateek, et al.
Published: (2025)
by: Chanda, Prateek, et al.
Published: (2025)
Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
by: Hu, Xinyue, et al.
Published: (2023)
by: Hu, Xinyue, et al.
Published: (2023)
VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
by: Bhope, Rahul Atul, et al.
Published: (2026)
by: Bhope, Rahul Atul, et al.
Published: (2026)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
by: Weng, Weixi, et al.
Published: (2024)
by: Weng, Weixi, et al.
Published: (2024)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
by: Akl, Ahmed, et al.
Published: (2024)
by: Akl, Ahmed, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning
by: Alavi, Ali
Published: (2026)
by: Alavi, Ali
Published: (2026)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
by: Huai, Tianyu, et al.
Published: (2025)
by: Huai, Tianyu, et al.
Published: (2025)
SCoRe: Submodular Combinatorial Representation Learning
by: Majee, Anay, et al.
Published: (2023)
by: Majee, Anay, et al.
Published: (2023)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
by: Indrehus, Kjetil, et al.
Published: (2026)
by: Indrehus, Kjetil, et al.
Published: (2026)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
by: Xie, Stephan, et al.
Published: (2026)
by: Xie, Stephan, et al.
Published: (2026)
SPDMark: Selective Parameter Displacement for Robust Video Watermarking
by: Fares, Samar, et al.
Published: (2025)
by: Fares, Samar, et al.
Published: (2025)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
by: Chen, Peiyuan, et al.
Published: (2024)
by: Chen, Peiyuan, et al.
Published: (2024)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
by: Chahe, Amirhosein, et al.
Published: (2025)
by: Chahe, Amirhosein, et al.
Published: (2025)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
by: Hartsock, Iryna, et al.
Published: (2024)
by: Hartsock, Iryna, et al.
Published: (2024)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
by: Li, Bingxin
Published: (2025)
by: Li, Bingxin
Published: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
by: Sun, Guohao, et al.
Published: (2024)
by: Sun, Guohao, et al.
Published: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
by: Srivastava, Varun, et al.
Published: (2025)
by: Srivastava, Varun, et al.
Published: (2025)
Answering Questions in Stages: Prompt Chaining for Contract QA
by: Roegiest, Adam, et al.
Published: (2024)
by: Roegiest, Adam, et al.
Published: (2024)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025)
by: Das, Swadhin, et al.
Published: (2025)
Similar Items
-
Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
by: Gupta, Madhav, et al.
Published: (2025) -
Less is More: Fewer Interpretable Region via Submodular Subset Selection
by: Chen, Ruoyu, et al.
Published: (2024) -
Improving Video Question Answering through query-based frame selection
by: Patil, Himanshu, et al.
Published: (2026) -
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
by: Chintapatla, Ishant, et al.
Published: (2025) -
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)