M3T: Multi-Modal Medical Transformer to bridge Clinical Context with Visual Insights for Retinal Image Medical Description Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Shaik, Nagur Shareef, Cherukuri, Teja Krishna, Ye, Dong Hye |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation
by: Shaik, Nagur Shareef, et al.
Published: (2026)
by: Shaik, Nagur Shareef, et al.
Published: (2026)
GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
Guided Context Gating: Learning to leverage salient lesions in retinal fundus images
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
Region-Affinity Attention for Whole-Slide Breast Cancer Classification in Deep Ultraviolet Imaging
by: Shaik, Nagur Shareef, et al.
Published: (2026)
by: Shaik, Nagur Shareef, et al.
Published: (2026)
Multi-modal Imaging Genomics Transformer: Attentive Integration of Imaging with Genomic Biomarkers for Schizophrenia Classification
by: Shaik, Nagur Shareef, et al.
Published: (2024)
by: Shaik, Nagur Shareef, et al.
Published: (2024)
DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
by: Shaik, Nagur Shareef, et al.
Published: (2025)
by: Shaik, Nagur Shareef, et al.
Published: (2025)
Spatial Sequence Attention Network for Schizophrenia Classification from Structural Brain MR Images
by: Shaik, Nagur Shareef, et al.
Published: (2024)
by: Shaik, Nagur Shareef, et al.
Published: (2024)
Ordinal Label-Distribution Learning with Constrained Asymmetric Priors for Imbalanced Retinal Grading
by: Shaik, Nagur Shareef, et al.
Published: (2025)
by: Shaik, Nagur Shareef, et al.
Published: (2025)
Dynamic Contextual Attention Network: Transforming Spatial Representations into Adaptive Insights for Endoscopic Polyp Diagnosis
by: Cherukuri, Teja Krishna, et al.
Published: (2025)
by: Cherukuri, Teja Krishna, et al.
Published: (2025)
DeepEyeNet: Generating Medical Report for Retinal Images
by: Huang, Jia-Hong
Published: (2025)
by: Huang, Jia-Hong
Published: (2025)
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
by: Ge, Hongyu, et al.
Published: (2025)
by: Ge, Hongyu, et al.
Published: (2025)
Medical Image Registration and Its Application in Retinal Images: A Review
by: Nie, Qiushi, et al.
Published: (2024)
by: Nie, Qiushi, et al.
Published: (2024)
Merging Context Clustering with Visual State Space Models for Medical Image Segmentation
by: Zhu, Yun, et al.
Published: (2025)
by: Zhu, Yun, et al.
Published: (2025)
Test-Time Modality Generalization for Medical Image Segmentation
by: Nam, Ju-Hyeon, et al.
Published: (2025)
by: Nam, Ju-Hyeon, et al.
Published: (2025)
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
by: Bai, Fan, et al.
Published: (2024)
by: Bai, Fan, et al.
Published: (2024)
Med3DInsight: Enhancing 3D Medical Image Understanding with 2D Multi-Modal Large Language Models
by: Chen, Qiuhui, et al.
Published: (2024)
by: Chen, Qiuhui, et al.
Published: (2024)
Automated Retinal Image Analysis and Medical Report Generation through Deep Learning
by: Huang, Jia-Hong
Published: (2024)
by: Huang, Jia-Hong
Published: (2024)
MedGS: Gaussian Splatting for Multi-Modal 3D Medical Imaging
by: Marzol, Kacper, et al.
Published: (2025)
by: Marzol, Kacper, et al.
Published: (2025)
Modality-Agnostic Structural Image Representation Learning for Deformable Multi-Modality Medical Image Registration
by: Mok, Tony C. W., et al.
Published: (2024)
by: Mok, Tony C. W., et al.
Published: (2024)
Comparing Image Segmentation Algorithms
by: Cherukuri, Milind
Published: (2025)
by: Cherukuri, Milind
Published: (2025)
Universal Vessel Segmentation for Multi-Modality Retinal Images
by: Wen, Bo, et al.
Published: (2025)
by: Wen, Bo, et al.
Published: (2025)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging
by: Pan, Tan, et al.
Published: (2026)
by: Pan, Tan, et al.
Published: (2026)
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
by: Agarwal, Lakshita, et al.
Published: (2025)
by: Agarwal, Lakshita, et al.
Published: (2025)
Context-driven Missing-Modality Learning for Robust Medical Diagnosis with Image-Tabular Data
by: Liu, Tianling, et al.
Published: (2026)
by: Liu, Tianling, et al.
Published: (2026)
Cross Modality Image Translation In Medical Imaging Using Generative Frameworks
by: Romoli, Giulia, et al.
Published: (2026)
by: Romoli, Giulia, et al.
Published: (2026)
Semi-Supervised Multi-Modal Medical Image Segmentation for Complex Situations
by: Meng, Dongdong, et al.
Published: (2025)
by: Meng, Dongdong, et al.
Published: (2025)
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
by: Xu, Ronghao, et al.
Published: (2025)
by: Xu, Ronghao, et al.
Published: (2025)
Mono-Modalizing Extremely Heterogeneous Multi-Modal Medical Image Registration
by: Choo, Kyobin, et al.
Published: (2025)
by: Choo, Kyobin, et al.
Published: (2025)
DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation
by: Lan, Libin, et al.
Published: (2025)
by: Lan, Libin, et al.
Published: (2025)
MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant
by: Zhan, Chenlu, et al.
Published: (2024)
by: Zhan, Chenlu, et al.
Published: (2024)
Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
Towards Robust In-Context Learning for Medical Image Segmentation via Data Synthesis
by: Hu, Jiesi, et al.
Published: (2025)
by: Hu, Jiesi, et al.
Published: (2025)
Multi-Conditioned Denoising Diffusion Probabilistic Model (mDDPM) for Medical Image Synthesis
by: Krishna, Arjun, et al.
Published: (2024)
by: Krishna, Arjun, et al.
Published: (2024)
Cycle Context Verification for In-Context Medical Image Segmentation
by: Hu, Shishuai, et al.
Published: (2025)
by: Hu, Shishuai, et al.
Published: (2025)
Efficient Parameter Adaptation for Multi-Modal Medical Image Segmentation and Prognosis
by: Saeed, Numan, et al.
Published: (2025)
by: Saeed, Numan, et al.
Published: (2025)
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
by: Huang, Peng, et al.
Published: (2025)
by: Huang, Peng, et al.
Published: (2025)
Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Global and Local Mamba Network for Multi-Modality Medical Image Super-Resolution
by: Ji, Zexin, et al.
Published: (2025)
by: Ji, Zexin, et al.
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
Similar Items
-
DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation
by: Shaik, Nagur Shareef, et al.
Published: (2026) -
GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning
by: Cherukuri, Teja Krishna, et al.
Published: (2024) -
Guided Context Gating: Learning to leverage salient lesions in retinal fundus images
by: Cherukuri, Teja Krishna, et al.
Published: (2024) -
Region-Affinity Attention for Whole-Slide Breast Cancer Classification in Deep Ultraviolet Imaging
by: Shaik, Nagur Shareef, et al.
Published: (2026) -
Multi-modal Imaging Genomics Transformer: Attentive Integration of Imaging with Genomic Biomarkers for Schizophrenia Classification
by: Shaik, Nagur Shareef, et al.
Published: (2024)