Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914126272397312 |
|---|---|
| author | Suresh, Yogesh Thakku Hogale, Vishwajeet Shivaji Zamfira, Luca-Alexandru Hegde, Anandavardhana |
| author_facet | Suresh, Yogesh Thakku Hogale, Vishwajeet Shivaji Zamfira, Luca-Alexandru Hegde, Anandavardhana |
| contents | We present a transformer-based multimodal framework for generating clinically relevant captions for MRI scans. Our system combines a DEiT-Small vision transformer as an image encoder, MediCareBERT for caption embedding, and a custom LSTM-based decoder. The architecture is designed to semantically align image and textual embeddings, using hybrid cosine-MSE loss and contrastive inference via vector similarity. We benchmark our method on the MultiCaRe dataset, comparing performance on filtered brain-only MRIs versus general MRI images against state-of-the-art medical image captioning methods including BLIP, R2GenGPT, and recent transformer-based approaches. Results show that focusing on domain-specific data improves caption accuracy and semantic alignment. Our work proposes a scalable, interpretable solution for automated medical image reporting. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25164 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning Suresh, Yogesh Thakku Hogale, Vishwajeet Shivaji Zamfira, Luca-Alexandru Hegde, Anandavardhana Image and Video Processing Artificial Intelligence Computer Vision and Pattern Recognition We present a transformer-based multimodal framework for generating clinically relevant captions for MRI scans. Our system combines a DEiT-Small vision transformer as an image encoder, MediCareBERT for caption embedding, and a custom LSTM-based decoder. The architecture is designed to semantically align image and textual embeddings, using hybrid cosine-MSE loss and contrastive inference via vector similarity. We benchmark our method on the MultiCaRe dataset, comparing performance on filtered brain-only MRIs versus general MRI images against state-of-the-art medical image captioning methods including BLIP, R2GenGPT, and recent transformer-based approaches. Results show that focusing on domain-specific data improves caption accuracy and semantic alignment. Our work proposes a scalable, interpretable solution for automated medical image reporting. |
| title | Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning |
| topic | Image and Video Processing Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.25164 |