Towards Domain Specification of Embedding Models in Medicine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khodadad, Mohammad, Kasmaee, Ali Shiraee, Astaraki, Mahdi, Mahyar, Hamidreza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913977548668928
author Khodadad, Mohammad
Kasmaee, Ali Shiraee
Astaraki, Mahdi
Mahyar, Hamidreza
author_facet Khodadad, Mohammad
Kasmaee, Ali Shiraee
Astaraki, Mahdi
Mahyar, Hamidreza
contents Medical text embedding models are foundational to a wide array of healthcare applications, ranging from clinical decision support and biomedical information retrieval to medical question answering, yet they remain hampered by two critical shortcomings. First, most models are trained on a narrow slice of medical and biological data, beside not being up to date in terms of methodology, making them ill suited to capture the diversity of terminology and semantics encountered in practice. Second, existing evaluations are often inadequate: even widely used benchmarks fail to generalize across the full spectrum of real world medical tasks. To address these gaps, we leverage MEDTE, a GTE model extensively fine-tuned on diverse medical corpora through self-supervised contrastive learning across multiple data sources, to deliver robust medical text embeddings. Alongside this model, we propose a comprehensive benchmark suite of 51 tasks spanning classification, clustering, pair classification, and retrieval modeled on the Massive Text Embedding Benchmark (MTEB) but tailored to the nuances of medical text. Our results demonstrate that this combined approach not only establishes a robust evaluation framework but also yields embeddings that consistently outperform state of the art alternatives in different tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19407
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Domain Specification of Embedding Models in Medicine
Khodadad, Mohammad
Kasmaee, Ali Shiraee
Astaraki, Mahdi
Mahyar, Hamidreza
Computation and Language
Medical text embedding models are foundational to a wide array of healthcare applications, ranging from clinical decision support and biomedical information retrieval to medical question answering, yet they remain hampered by two critical shortcomings. First, most models are trained on a narrow slice of medical and biological data, beside not being up to date in terms of methodology, making them ill suited to capture the diversity of terminology and semantics encountered in practice. Second, existing evaluations are often inadequate: even widely used benchmarks fail to generalize across the full spectrum of real world medical tasks. To address these gaps, we leverage MEDTE, a GTE model extensively fine-tuned on diverse medical corpora through self-supervised contrastive learning across multiple data sources, to deliver robust medical text embeddings. Alongside this model, we propose a comprehensive benchmark suite of 51 tasks spanning classification, clustering, pair classification, and retrieval modeled on the Massive Text Embedding Benchmark (MTEB) but tailored to the nuances of medical text. Our results demonstrate that this combined approach not only establishes a robust evaluation framework but also yields embeddings that consistently outperform state of the art alternatives in different tasks.
title Towards Domain Specification of Embedding Models in Medicine
topic Computation and Language
url https://arxiv.org/abs/2507.19407