Improving Medical VQA through Trajectory-Aware Process Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Gulluk, Halil Ibrahim, Gevaert, Olivier |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026)
by: Byun, Ji Young, et al.
Published: (2026)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
by: Ma, Yongpei, et al.
Published: (2025)
by: Ma, Yongpei, et al.
Published: (2025)
Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization
by: Wang, Jiayun, et al.
Published: (2024)
by: Wang, Jiayun, et al.
Published: (2024)
CAD: Confidence-Aware Adaptive Displacement for Semi-Supervised Medical Image Segmentation
by: Xiao, Wenbo, et al.
Published: (2025)
by: Xiao, Wenbo, et al.
Published: (2025)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
by: Nguyen, Khoi Anh, et al.
Published: (2025)
by: Nguyen, Khoi Anh, et al.
Published: (2025)
Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects
by: Warner, Elisa, et al.
Published: (2023)
by: Warner, Elisa, et al.
Published: (2023)
SITUATE: Indoor Human Trajectory Prediction through Geometric Features and Self-Supervised Vision Representation
by: Capogrosso, Luigi, et al.
Published: (2024)
by: Capogrosso, Luigi, et al.
Published: (2024)
WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring
by: Habibpour, Mobin, et al.
Published: (2026)
by: Habibpour, Mobin, et al.
Published: (2026)
CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
by: Ding, Tianxingjian, et al.
Published: (2025)
by: Ding, Tianxingjian, et al.
Published: (2025)
Disentanglement-Based Equivariant Learning for Compositional VQA
by: Du, Zhou, et al.
Published: (2026)
by: Du, Zhou, et al.
Published: (2026)
Unexplored flaws in multiple-choice VQA evaluations
by: Rosenthal, Fabio, et al.
Published: (2025)
by: Rosenthal, Fabio, et al.
Published: (2025)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
RNNs, CNNs and Transformers in Human Action Recognition: A Survey and a Hybrid Model
by: Alomar, Khaled, et al.
Published: (2024)
by: Alomar, Khaled, et al.
Published: (2024)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
Synergy-Guided Regional Supervision of Pseudo Labels for Semi-Supervised Medical Image Segmentation
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
Improving Colorectal Cancer Screening and Risk Assessment through Predictive Modeling on Medical Images and Records
by: Jiang, Shuai, et al.
Published: (2024)
by: Jiang, Shuai, et al.
Published: (2024)
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024)
by: Baitieva, Aimira, et al.
Published: (2024)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision
by: Sun, Tianyao, et al.
Published: (2025)
by: Sun, Tianyao, et al.
Published: (2025)
Learning Velocity and Acceleration: Self-Supervised Motion Consistency for Pedestrian Trajectory Prediction
by: Huang, Yizhou, et al.
Published: (2025)
by: Huang, Yizhou, et al.
Published: (2025)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
HAMMR: HierArchical MultiModal React agents for generic VQA
by: Castrejon, Lluis, et al.
Published: (2024)
by: Castrejon, Lluis, et al.
Published: (2024)
AI-Derived Structural Building Intelligence for Urban Resilience: An Application in Saint Vincent and the Grenadines
by: Tingzon, Isabelle, et al.
Published: (2025)
by: Tingzon, Isabelle, et al.
Published: (2025)
Beyond Conventional Transformers: The Medical X-ray Attention (MXA) Block for Improved Multi-Label Diagnosis Using Knowledge Distillation
by: Rand, Amit, et al.
Published: (2025)
by: Rand, Amit, et al.
Published: (2025)
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023)
by: Berlot-Attwell, Ian, et al.
Published: (2023)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
by: Zhou, Runjie, et al.
Published: (2026)
by: Zhou, Runjie, et al.
Published: (2026)
Evaluating Feature Attribution Methods in the Image Domain
by: Gevaert, Arne, et al.
Published: (2022)
by: Gevaert, Arne, et al.
Published: (2022)
Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
by: Cao, Ji, et al.
Published: (2025)
by: Cao, Ji, et al.
Published: (2025)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
by: Gong, Yunye, et al.
Published: (2023)
by: Gong, Yunye, et al.
Published: (2023)
Similarity Trajectories: Linking Sampling Process to Artifacts in Diffusion-Generated Images
by: Menn, Dennis, et al.
Published: (2024)
by: Menn, Dennis, et al.
Published: (2024)
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
by: Das, Aryan, et al.
Published: (2026)
by: Das, Aryan, et al.
Published: (2026)
Diffusion-Based Environment-Aware Trajectory Prediction
by: Westny, Theodor, et al.
Published: (2024)
by: Westny, Theodor, et al.
Published: (2024)
On Improving the Algorithm-, Model-, and Data- Efficiency of Self-Supervised Learning
by: Cao, Yun-Hao, et al.
Published: (2024)
by: Cao, Yun-Hao, et al.
Published: (2024)
Pseudo-label Refinement for Improving Self-Supervised Learning Systems
by: Zia-ur-Rehman, et al.
Published: (2024)
by: Zia-ur-Rehman, et al.
Published: (2024)
Can Generative Models Improve Self-Supervised Representation Learning?
by: Ayromlou, Sana, et al.
Published: (2024)
by: Ayromlou, Sana, et al.
Published: (2024)
Similar Items
-
MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification
by: Gulluk, Halil Ibrahim, et al.
Published: (2026) -
SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning
by: Gulluk, Halil Ibrahim, et al.
Published: (2026) -
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026) -
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024) -
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
by: Ma, Yongpei, et al.
Published: (2025)