Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yongpei, Wang, Pengyu, Dunn, Adam, Naseem, Usman, Kim, Jinman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation
by: Ye, Shuchang, et al.
Published: (2025)
by: Ye, Shuchang, et al.
Published: (2025)
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
by: Ye, Shuchang, et al.
Published: (2025)
by: Ye, Shuchang, et al.
Published: (2025)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
Improving Medical VQA through Trajectory-Aware Process Supervision
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
Dynamic Traceback Learning for Medical Report Generation
by: Ye, Shuchang, et al.
Published: (2024)
by: Ye, Shuchang, et al.
Published: (2024)
Bridging the Gap: Protocol Towards Fair and Consistent Affect Analysis
by: Hu, Guanyu, et al.
Published: (2024)
by: Hu, Guanyu, et al.
Published: (2024)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
by: Mo, Wentao, et al.
Published: (2024)
by: Mo, Wentao, et al.
Published: (2024)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
by: Vaish, Puru, et al.
Published: (2024)
by: Vaish, Puru, et al.
Published: (2024)
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023)
by: Berlot-Attwell, Ian, et al.
Published: (2023)
Improving Consistency Models with Generator-Augmented Flows
by: Issenhuth, Thibaut, et al.
Published: (2024)
by: Issenhuth, Thibaut, et al.
Published: (2024)
Transfer or Self-Supervised? Bridging the Performance Gap in Medical Imaging
by: Zhao, Zehui, et al.
Published: (2024)
by: Zhao, Zehui, et al.
Published: (2024)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026)
by: Byun, Ji Young, et al.
Published: (2026)
RadGenome-Anatomy: A Large-Scale Anatomy-Labeled Chest Radiograph Dataset via Physically Grounded Volumetric Projection
by: Ye, Shuchang, et al.
Published: (2026)
by: Ye, Shuchang, et al.
Published: (2026)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
by: Nguyen, Khoi Anh, et al.
Published: (2025)
by: Nguyen, Khoi Anh, et al.
Published: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
by: Van-Dinh, Tue-Thu, et al.
Published: (2025)
Consistency Diffusion Bridge Models
by: He, Guande, et al.
Published: (2024)
by: He, Guande, et al.
Published: (2024)
Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation
by: Zhang, Zhe, et al.
Published: (2026)
by: Zhang, Zhe, et al.
Published: (2026)
Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
by: Vaish, Puru, et al.
Published: (2025)
by: Vaish, Puru, et al.
Published: (2025)
Informed Mixing -- Improving Open Set Recognition via Attribution-based Augmentation
by: Xu, Jiawen, et al.
Published: (2025)
by: Xu, Jiawen, et al.
Published: (2025)
Semantically Consistent Video Inpainting with Conditional Diffusion Models
by: Green, Dylan, et al.
Published: (2024)
by: Green, Dylan, et al.
Published: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
ScribbleGen: Generative Data Augmentation Improves Scribble-supervised Semantic Segmentation
by: Schnell, Jacob, et al.
Published: (2023)
by: Schnell, Jacob, et al.
Published: (2023)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
by: Rao, Qihang, et al.
Published: (2025)
by: Rao, Qihang, et al.
Published: (2025)
WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring
by: Habibpour, Mobin, et al.
Published: (2026)
by: Habibpour, Mobin, et al.
Published: (2026)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
by: Diao, Xingjian, et al.
Published: (2025)
by: Diao, Xingjian, et al.
Published: (2025)
Stable Consistency Tuning: Understanding and Improving Consistency Models
by: Wang, Fu-Yun, et al.
Published: (2024)
by: Wang, Fu-Yun, et al.
Published: (2024)
Towards Understanding Why Data Augmentation Improves Generalization
by: Li, Jingyang, et al.
Published: (2025)
by: Li, Jingyang, et al.
Published: (2025)
Disentanglement-Based Equivariant Learning for Compositional VQA
by: Du, Zhou, et al.
Published: (2026)
by: Du, Zhou, et al.
Published: (2026)
Unexplored flaws in multiple-choice VQA evaluations
by: Rosenthal, Fabio, et al.
Published: (2025)
by: Rosenthal, Fabio, et al.
Published: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
TinyVQA: Compact Multimodal Deep Neural Network for Visual Question Answering on Resource-Constrained Devices
by: Rashid, Hasib-Al, et al.
Published: (2024)
by: Rashid, Hasib-Al, et al.
Published: (2024)
Closing the Modality Gap Aligns Group-Wise Semantics
by: Grassucci, Eleonora, et al.
Published: (2026)
by: Grassucci, Eleonora, et al.
Published: (2026)
Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift
by: Khan, Behraj, et al.
Published: (2025)
by: Khan, Behraj, et al.
Published: (2025)
Taking Shortcuts for Categorical VQA Using Super Neurons
by: Musacchio, Pierre, et al.
Published: (2026)
by: Musacchio, Pierre, et al.
Published: (2026)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
by: Ma, Haodi, et al.
Published: (2025)
by: Ma, Haodi, et al.
Published: (2025)
Age-Diverse Deepfake Dataset: Bridging the Age Gap in Deepfake Detection
by: Joshi, Unisha
Published: (2025)
by: Joshi, Unisha
Published: (2025)
Similar Items
-
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025) -
Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation
by: Ye, Shuchang, et al.
Published: (2025) -
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
by: Ye, Shuchang, et al.
Published: (2025) -
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024) -
Improving Medical VQA through Trajectory-Aware Process Supervision
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)