Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Suresh, Yogesh Thakku, Hogale, Vishwajeet Shivaji, Zamfira, Luca-Alexandru, Hegde, Anandavardhana
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914126272397312
author Suresh, Yogesh Thakku
Hogale, Vishwajeet Shivaji
Zamfira, Luca-Alexandru
Hegde, Anandavardhana
author_facet Suresh, Yogesh Thakku
Hogale, Vishwajeet Shivaji
Zamfira, Luca-Alexandru
Hegde, Anandavardhana
contents We present a transformer-based multimodal framework for generating clinically relevant captions for MRI scans. Our system combines a DEiT-Small vision transformer as an image encoder, MediCareBERT for caption embedding, and a custom LSTM-based decoder. The architecture is designed to semantically align image and textual embeddings, using hybrid cosine-MSE loss and contrastive inference via vector similarity. We benchmark our method on the MultiCaRe dataset, comparing performance on filtered brain-only MRIs versus general MRI images against state-of-the-art medical image captioning methods including BLIP, R2GenGPT, and recent transformer-based approaches. Results show that focusing on domain-specific data improves caption accuracy and semantic alignment. Our work proposes a scalable, interpretable solution for automated medical image reporting.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning
Suresh, Yogesh Thakku
Hogale, Vishwajeet Shivaji
Zamfira, Luca-Alexandru
Hegde, Anandavardhana
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
We present a transformer-based multimodal framework for generating clinically relevant captions for MRI scans. Our system combines a DEiT-Small vision transformer as an image encoder, MediCareBERT for caption embedding, and a custom LSTM-based decoder. The architecture is designed to semantically align image and textual embeddings, using hybrid cosine-MSE loss and contrastive inference via vector similarity. We benchmark our method on the MultiCaRe dataset, comparing performance on filtered brain-only MRIs versus general MRI images against state-of-the-art medical image captioning methods including BLIP, R2GenGPT, and recent transformer-based approaches. Results show that focusing on domain-specific data improves caption accuracy and semantic alignment. Our work proposes a scalable, interpretable solution for automated medical image reporting.
title Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.25164