Saved in:
| Main Authors: | Hagen, Luca, Müller, Johanna P., Zhang, Weitong, Qiao, Mengyun, Kainz, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.18313 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Resource-efficient Medical Image Analysis with Self-adapting Forward-Forward Networks
by: Müller, Johanna P., et al.
Published: (2024)
by: Müller, Johanna P., et al.
Published: (2024)
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
by: Li, Zuoou, et al.
Published: (2025)
by: Li, Zuoou, et al.
Published: (2025)
Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
by: Erick, Franciskus Xaverius, et al.
Published: (2026)
by: Erick, Franciskus Xaverius, et al.
Published: (2026)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
by: Xi, Suyang, et al.
Published: (2026)
by: Xi, Suyang, et al.
Published: (2026)
Free Form Medical Visual Question Answering in Radiology
by: Narayanan, Abhishek, et al.
Published: (2024)
by: Narayanan, Abhishek, et al.
Published: (2024)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Hallucination Benchmark in Medical Visual Question Answering
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
by: Rajkhowa, Tonmoy, et al.
Published: (2024)
by: Rajkhowa, Tonmoy, et al.
Published: (2024)
Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging
by: Kainz, Bernhard, et al.
Published: (2026)
by: Kainz, Bernhard, et al.
Published: (2026)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
Graph Conditioned Diffusion for Controllable Histopathology Image Generation
by: Cechnicka, Sarah, et al.
Published: (2025)
by: Cechnicka, Sarah, et al.
Published: (2025)
Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases
by: Zhu, Huanjia, et al.
Published: (2025)
by: Zhu, Huanjia, et al.
Published: (2025)
Uncovering Hidden Subspaces in Video Diffusion Models Using Re-Identification
by: Dombrowski, Mischa, et al.
Published: (2024)
by: Dombrowski, Mischa, et al.
Published: (2024)
Explicit Abstention Knobs for Predictable Reliability in Video Question Answering
by: Ortiz, Jorge
Published: (2025)
by: Ortiz, Jorge
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering
by: Zhang, Zhilin, et al.
Published: (2024)
by: Zhang, Zhilin, et al.
Published: (2024)
VoQA: Visual-only Question Answering
by: An, Jianing, et al.
Published: (2025)
by: An, Jianing, et al.
Published: (2025)
Object Attribute Matters in Visual Question Answering
by: Li, Peize, et al.
Published: (2023)
by: Li, Peize, et al.
Published: (2023)
Diffusing the Blind Spot: Uterine MRI Synthesis with Diffusion Models
by: Müller, Johanna P., et al.
Published: (2025)
by: Müller, Johanna P., et al.
Published: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
by: Zhang, Zhengxuan, et al.
Published: (2025)
by: Zhang, Zhengxuan, et al.
Published: (2025)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
by: Ahir, Param, et al.
Published: (2023)
by: Ahir, Param, et al.
Published: (2023)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
by: Liu, Ying, et al.
Published: (2024)
by: Liu, Ying, et al.
Published: (2024)
Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
by: Marouf, Imad Eddine, et al.
Published: (2025)
by: Marouf, Imad Eddine, et al.
Published: (2025)
LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
by: Ma, Runze, et al.
Published: (2026)
by: Ma, Runze, et al.
Published: (2026)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
by: Liu, Zhiyue, et al.
Published: (2025)
by: Liu, Zhiyue, et al.
Published: (2025)
Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering
by: Felizzi, Federico, et al.
Published: (2025)
by: Felizzi, Federico, et al.
Published: (2025)
Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
by: Li, Zhifei, et al.
Published: (2025)
by: Li, Zhifei, et al.
Published: (2025)
StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
by: Wen, Zhihao, et al.
Published: (2025)
by: Wen, Zhihao, et al.
Published: (2025)
LingoQA: Visual Question Answering for Autonomous Driving
by: Marcu, Ana-Maria, et al.
Published: (2023)
by: Marcu, Ana-Maria, et al.
Published: (2023)
CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
by: Vu, Sinh Trong, et al.
Published: (2025)
by: Vu, Sinh Trong, et al.
Published: (2025)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
by: Jain, Riddhi, et al.
Published: (2025)
by: Jain, Riddhi, et al.
Published: (2025)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
Similar Items
-
Resource-efficient Medical Image Analysis with Self-adapting Forward-Forward Networks
by: Müller, Johanna P., et al.
Published: (2024) -
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
by: Li, Zuoou, et al.
Published: (2025) -
Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
by: Erick, Franciskus Xaverius, et al.
Published: (2026) -
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
by: Xi, Suyang, et al.
Published: (2026) -
Free Form Medical Visual Question Answering in Radiology
by: Narayanan, Abhishek, et al.
Published: (2024)