Medico 2025: Visual Question Answering for Gastrointestinal Imaging
Fuente:
arXiv
Saved in:
| Main Authors: | Gautam, Sushant, Thambawita, Vajira, Riegler, Michael, Halvorsen, Pål, Hicks, Steven |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
by: Ma, Siteng, et al.
Published: (2025)
by: Ma, Siteng, et al.
Published: (2025)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
by: Chaichuk, Mikhail, et al.
Published: (2025)
by: Chaichuk, Mikhail, et al.
Published: (2025)
SelectiveKD: A semi-supervised framework for cancer detection in DBT through Knowledge Distillation and Pseudo-labeling
by: Dillard, Laurent, et al.
Published: (2024)
by: Dillard, Laurent, et al.
Published: (2024)
Automated Cervical Cancer Detection through Visual Inspection with Acetic Acid in Resource-Poor Settings with Lightweight Deep Learning Models Deployed on an Android Device
by: Maben, Leander Melroy, et al.
Published: (2025)
by: Maben, Leander Melroy, et al.
Published: (2025)
DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
by: Çalışkan, Halil Hüseyin, et al.
Published: (2025)
by: Çalışkan, Halil Hüseyin, et al.
Published: (2025)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
by: Khan, Md Ashik, et al.
Published: (2025)
by: Khan, Md Ashik, et al.
Published: (2025)
MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models
by: Balezo, Guillaume, et al.
Published: (2025)
by: Balezo, Guillaume, et al.
Published: (2025)
GAN-GA: A Generative Model based on Genetic Algorithm for Medical Image Generation
by: AbdulRazek, M., et al.
Published: (2023)
by: AbdulRazek, M., et al.
Published: (2023)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
by: Long, Yuchong, et al.
Published: (2025)
by: Long, Yuchong, et al.
Published: (2025)
Uncertainty-Calibrated Explainable Artificial Intelligence for Fetal Ultrasound Plane Classification: A Systematic Review
by: Lundström-Imanov, Gustav Olaf Yunus Laitinen-Fredriksson, et al.
Published: (2026)
by: Lundström-Imanov, Gustav Olaf Yunus Laitinen-Fredriksson, et al.
Published: (2026)
SLIM-Diff: Shared Latent Image-Mask Diffusion with Lp loss for Data-Scarce Epilepsy FLAIR MRI
by: Pascual-González, Mario, et al.
Published: (2026)
by: Pascual-González, Mario, et al.
Published: (2026)
Balanced conic rectified flow
by: Kim, Shin Seong, et al.
Published: (2025)
by: Kim, Shin Seong, et al.
Published: (2025)
ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection
by: Samson, Hema Hariharan
Published: (2026)
by: Samson, Hema Hariharan
Published: (2026)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
Retinal Fundus Multi-Disease Image Classification using Hybrid CNN-Transformer-Ensemble Architectures
by: Singh, Deependra, et al.
Published: (2025)
by: Singh, Deependra, et al.
Published: (2025)
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines
by: Wimalasiri, Chathura
Published: (2026)
by: Wimalasiri, Chathura
Published: (2026)
Rethinking VLMs for Image Forgery Detection and Localization
by: Guo, Shaofeng, et al.
Published: (2026)
by: Guo, Shaofeng, et al.
Published: (2026)
S3Simulator: A benchmarking Side Scan Sonar Simulator dataset for Underwater Image Analysis
by: S, Kamal Basha, et al.
Published: (2024)
by: S, Kamal Basha, et al.
Published: (2024)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
BreastDCEDL: A Comprehensive Breast Cancer DCE-MRI Dataset and Transformer Implementation for Treatment Response Prediction
by: Fridman, Naomi, et al.
Published: (2025)
by: Fridman, Naomi, et al.
Published: (2025)
Synthetic-to-Real Transfer Learning for Chromatin-Sensitive PWS Microscopy
by: Arafat, Jahidul, et al.
Published: (2025)
by: Arafat, Jahidul, et al.
Published: (2025)
Revealing an Unattractivity Bias in Mental Reconstruction of Occluded Faces using Generative Image Models
by: Riedmann, Frederik, et al.
Published: (2024)
by: Riedmann, Frederik, et al.
Published: (2024)
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
by: Ko, Hanbin, et al.
Published: (2025)
by: Ko, Hanbin, et al.
Published: (2025)
A Novel Approach to Breast Cancer Segmentation using U-Net Model with Attention Mechanisms and FedProx
by: Gad, Eyad, et al.
Published: (2025)
by: Gad, Eyad, et al.
Published: (2025)
Deep Learning for Generating Computational PIN-4 Immunohistochemistry Staining from Prostate Biopsy H&E Images
by: Tran, Vietbao, et al.
Published: (2026)
by: Tran, Vietbao, et al.
Published: (2026)
Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images
by: Da, Longchao, et al.
Published: (2024)
by: Da, Longchao, et al.
Published: (2024)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023)
by: Tourani, Ali, et al.
Published: (2023)
Technical Report: Automated Optical Inspection of Surgical Instruments
by: Shafqat, Zunaira, et al.
Published: (2026)
by: Shafqat, Zunaira, et al.
Published: (2026)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
by: Kalušev, Vladimir, et al.
Published: (2026)
by: Kalušev, Vladimir, et al.
Published: (2026)
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Person detection and re-identification in open-world settings of retail stores and public spaces
by: Brkljač, Branko, et al.
Published: (2025)
by: Brkljač, Branko, et al.
Published: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
A Nerf-Based Color Consistency Method for Remote Sensing Images
by: Zuo, Zongcheng, et al.
Published: (2024)
by: Zuo, Zongcheng, et al.
Published: (2024)
TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models
by: Medeiros, Daniel Nobrega
Published: (2026)
by: Medeiros, Daniel Nobrega
Published: (2026)
Blink-to-code: real-time Morse code communication via eye blink detection and classification
by: Bhatt, Anushka
Published: (2025)
by: Bhatt, Anushka
Published: (2025)
Similar Items
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
by: Gautam, Sushant, et al.
Published: (2025) -
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025) -
Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
by: Ma, Siteng, et al.
Published: (2025) -
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)