Semantic Alignment of Unimodal Medical Text and Vision Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Di Folco, Maxime, Chan, Emily, Hasny, Marta, Bercea, Cosmin I., Schnabel, Julia A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable Representation Learning of Cardiac MRI via Attribute Regularization
by: Di Folco, Maxime, et al.
Published: (2024)
by: Di Folco, Maxime, et al.
Published: (2024)
Tables Guide Vision: Learning to See the Heart through Tabular Data
by: Hasny, Marta, et al.
Published: (2025)
by: Hasny, Marta, et al.
Published: (2025)
No Data? No Problem: Robust Vision-Tabular Learning with Missing Values
by: Hasny, Marta, et al.
Published: (2025)
by: Hasny, Marta, et al.
Published: (2025)
Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification
by: Mayr, Josef, et al.
Published: (2025)
by: Mayr, Josef, et al.
Published: (2025)
Denoising Diffusion Models for Anomaly Localization in Medical Images
by: Bercea, Cosmin I., et al.
Published: (2024)
by: Bercea, Cosmin I., et al.
Published: (2024)
Towards Universal Unsupervised Anomaly Detection in Medical Imaging
by: Bercea, Cosmin I., et al.
Published: (2024)
by: Bercea, Cosmin I., et al.
Published: (2024)
Diffusion Models with Implicit Guidance for Medical Anomaly Detection
by: Bercea, Cosmin I., et al.
Published: (2024)
by: Bercea, Cosmin I., et al.
Published: (2024)
Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
LocBAM: Advancing 3D Patch-Based Image Segmentation by Integrating Location Contex
by: Hooft, Donnate, et al.
Published: (2026)
by: Hooft, Donnate, et al.
Published: (2026)
On Differentially Private 3D Medical Image Synthesis with Controllable Latent Diffusion Models
by: Daum, Deniz, et al.
Published: (2024)
by: Daum, Deniz, et al.
Published: (2024)
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
Unsupervised Analysis of Alzheimer's Disease Signatures using 3D Deformable Autoencoders
by: Avci, Mehmet Yigit, et al.
Published: (2024)
by: Avci, Mehmet Yigit, et al.
Published: (2024)
Influence of Classification Task and Distribution Shift Type on OOD Detection in Fetal Ultrasound
by: Wong, Chun Kit, et al.
Published: (2025)
by: Wong, Chun Kit, et al.
Published: (2025)
Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging
by: Kainz, Bernhard, et al.
Published: (2026)
by: Kainz, Bernhard, et al.
Published: (2026)
MedEdit: Counterfactual Diffusion-based Image Editing on Brain MRI
by: Alaya, Malek Ben, et al.
Published: (2024)
by: Alaya, Malek Ben, et al.
Published: (2024)
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks
by: Wu, Peiran, et al.
Published: (2024)
by: Wu, Peiran, et al.
Published: (2024)
Assessing and Learning Alignment of Unimodal Vision and Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
Language Models Meet Anomaly Detection for Better Interpretability and Generalizability
by: Li, Jun, et al.
Published: (2024)
by: Li, Jun, et al.
Published: (2024)
Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies
by: Schaper, Ben, et al.
Published: (2026)
by: Schaper, Ben, et al.
Published: (2026)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
by: Johnson, Emily, et al.
Published: (2025)
by: Johnson, Emily, et al.
Published: (2025)
General Vision Encoder Features as Guidance in Medical Image Registration
by: Kögl, Fryderyk, et al.
Published: (2024)
by: Kögl, Fryderyk, et al.
Published: (2024)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation
by: Maurya, Lalit, et al.
Published: (2025)
by: Maurya, Lalit, et al.
Published: (2025)
Data-Driven Tissue- and Subject-Specific Elastic Regularization for Medical Image Registration
by: Reithmeir, Anna, et al.
Published: (2024)
by: Reithmeir, Anna, et al.
Published: (2024)
Fast Context-Based Low-Light Image Enhancement via Neural Implicit Representations
by: Chobola, Tomáš, et al.
Published: (2024)
by: Chobola, Tomáš, et al.
Published: (2024)
Multi-Grained Cross-modal Alignment for Learning Open-vocabulary Semantic Segmentation from Text Supervision
by: Liu, Yajie, et al.
Published: (2024)
by: Liu, Yajie, et al.
Published: (2024)
Machine Unlearning for Medical Imaging
by: Nasirigerdeh, Reza, et al.
Published: (2024)
by: Nasirigerdeh, Reza, et al.
Published: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
by: Jiang, Jerry, et al.
Published: (2026)
by: Jiang, Jerry, et al.
Published: (2026)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
by: Shams, Montasir, et al.
Published: (2025)
by: Shams, Montasir, et al.
Published: (2025)
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
by: Jose, Cijo, et al.
Published: (2024)
by: Jose, Cijo, et al.
Published: (2024)
MedDIFT: Multi-Scale Diffusion-Based Correspondence in 3D Medical Imaging
by: Zhang, Xingyu, et al.
Published: (2025)
by: Zhang, Xingyu, et al.
Published: (2025)
Text-Guided Semantic Image Encoder
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
by: Liu, Che, et al.
Published: (2025)
by: Liu, Che, et al.
Published: (2025)
Improving Human Image Animation via Semantic Representation Alignment
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Similar Items
-
Interpretable Representation Learning of Cardiac MRI via Attribute Regularization
by: Di Folco, Maxime, et al.
Published: (2024) -
Tables Guide Vision: Learning to See the Heart through Tabular Data
by: Hasny, Marta, et al.
Published: (2025) -
No Data? No Problem: Robust Vision-Tabular Learning with Missing Values
by: Hasny, Marta, et al.
Published: (2025) -
Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification
by: Mayr, Josef, et al.
Published: (2025) -
Denoising Diffusion Models for Anomaly Localization in Medical Images
by: Bercea, Cosmin I., et al.
Published: (2024)