jina-embeddings-v5-text: Task-Targeted Embedding Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Akram, Mohammad Kalim, Sturua, Saba, Havriushenko, Nastia, Herreros, Quentin, Günther, Michael, Werk, Maximilian, Xiao, Han
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908998672842752
author Akram, Mohammad Kalim
Sturua, Saba
Havriushenko, Nastia
Herreros, Quentin
Günther, Michael
Werk, Maximilian
Xiao, Han
author_facet Akram, Mohammad Kalim
Sturua, Saba
Havriushenko, Nastia
Herreros, Quentin
Günther, Michael
Werk, Maximilian
Xiao, Han
contents Text embedding models are widely used for semantic similarity tasks, including information retrieval, clustering, and classification. General-purpose models are typically trained with single- or multi-stage processes using contrastive loss functions. We introduce a novel training regimen that combines model distillation techniques with task-specific contrastive loss to produce compact, high-performance embedding models. Our findings suggest that this approach is more effective for training small models than purely contrastive or distillation-based training paradigms alone. Benchmark scores for the resulting models, jina-embeddings-v5-text-small and jina-embeddings-v5-text-nano, exceed or match the state-of-the-art for models of similar size. jina-embeddings-v5-text models additionally support long texts (up to 32k tokens) in many languages, and generate embeddings that remain robust under truncation and binary quantization. Model weights are publicly available, hopefully inspiring further advances in embedding model development.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15547
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle jina-embeddings-v5-text: Task-Targeted Embedding Distillation
Akram, Mohammad Kalim
Sturua, Saba
Havriushenko, Nastia
Herreros, Quentin
Günther, Michael
Werk, Maximilian
Xiao, Han
Computation and Language
Text embedding models are widely used for semantic similarity tasks, including information retrieval, clustering, and classification. General-purpose models are typically trained with single- or multi-stage processes using contrastive loss functions. We introduce a novel training regimen that combines model distillation techniques with task-specific contrastive loss to produce compact, high-performance embedding models. Our findings suggest that this approach is more effective for training small models than purely contrastive or distillation-based training paradigms alone. Benchmark scores for the resulting models, jina-embeddings-v5-text-small and jina-embeddings-v5-text-nano, exceed or match the state-of-the-art for models of similar size. jina-embeddings-v5-text models additionally support long texts (up to 32k tokens) in many languages, and generate embeddings that remain robust under truncation and binary quantization. Model weights are publicly available, hopefully inspiring further advances in embedding model development.
title jina-embeddings-v5-text: Task-Targeted Embedding Distillation
topic Computation and Language
url https://arxiv.org/abs/2602.15547