Semantic Alignment of Unimodal Medical Text and Vision Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Di Folco, Maxime, Chan, Emily, Hasny, Marta, Bercea, Cosmin I., Schnabel, Julia A.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909527870275584
author Di Folco, Maxime
Chan, Emily
Hasny, Marta
Bercea, Cosmin I.
Schnabel, Julia A.
author_facet Di Folco, Maxime
Chan, Emily
Hasny, Marta
Bercea, Cosmin I.
Schnabel, Julia A.
contents General-purpose AI models, particularly those designed for text and vision, demonstrate impressive versatility across a wide range of deep-learning tasks. However, they often underperform in specialised domains like medical imaging, where domain-specific solutions or alternative knowledge transfer approaches are typically required. Recent studies have noted that general-purpose models can exhibit similar latent spaces when processing semantically related data, although this alignment does not occur naturally. Building on this insight, it has been shown that applying a simple transformation - at most affine - estimated from a subset of semantically corresponding samples, known as anchors, enables model stitching across diverse training paradigms, architectures, and modalities. In this paper, we explore how semantic alignment - estimating transformations between anchors - can bridge general-purpose AI with specialised medical knowledge. Using multiple public chest X-ray datasets, we demonstrate that model stitching across model architectures allows general models to integrate domain-specific knowledge without additional training, leading to improved performance on medical tasks. Furthermore, we introduce a novel zero-shot classification approach for unimodal vision encoders that leverages semantic alignment across modalities. Our results show that our method not only outperforms general multimodal models but also approaches the performance levels of fully trained, medical-specific multimodal solutions
format Preprint
id arxiv_https___arxiv_org_abs_2503_04478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic Alignment of Unimodal Medical Text and Vision Representations
Di Folco, Maxime
Chan, Emily
Hasny, Marta
Bercea, Cosmin I.
Schnabel, Julia A.
Computer Vision and Pattern Recognition
General-purpose AI models, particularly those designed for text and vision, demonstrate impressive versatility across a wide range of deep-learning tasks. However, they often underperform in specialised domains like medical imaging, where domain-specific solutions or alternative knowledge transfer approaches are typically required. Recent studies have noted that general-purpose models can exhibit similar latent spaces when processing semantically related data, although this alignment does not occur naturally. Building on this insight, it has been shown that applying a simple transformation - at most affine - estimated from a subset of semantically corresponding samples, known as anchors, enables model stitching across diverse training paradigms, architectures, and modalities. In this paper, we explore how semantic alignment - estimating transformations between anchors - can bridge general-purpose AI with specialised medical knowledge. Using multiple public chest X-ray datasets, we demonstrate that model stitching across model architectures allows general models to integrate domain-specific knowledge without additional training, leading to improved performance on medical tasks. Furthermore, we introduce a novel zero-shot classification approach for unimodal vision encoders that leverages semantic alignment across modalities. Our results show that our method not only outperforms general multimodal models but also approaches the performance levels of fully trained, medical-specific multimodal solutions
title Semantic Alignment of Unimodal Medical Text and Vision Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.04478