Saved in:
Bibliographic Details
Main Authors: Dhanalakshmi, Kommu, Sree, Dr. K. Santhi
Format: Recurso digital
Language:
Published: Zenodo 2025
Online Access:https://doi.org/10.5281/zenodo.15654851
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • <p>The work of Visual Question Answering (VQA) in the medical profession is challenging and involves both successfully answering field-specific questions and evaluating medical images. The deep learning-based approach for medical VQA is presented in this paper, which uses convolutional neural networks (CNNs) for image feature extraction and transformers for natural language processing. In order to improve reasoning over medical imagery and textual inquiries, the proposed model uses multi-modal embeddings. The evaluation dataset includes images of radiology and pathology along with corresponding question-answer pairs. The results demonstrate that our model performs better than existing methods, providing reliable and understandable responses to support medical diagnosis. We also explore explainability tactics to improve user trust and model transparency</p>