Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yanzhao, Li, Mingxin, Long, Dingkun, Zhang, Xin, Lin, Huan, Yang, Baosong, Xie, Pengjun, Yang, An, Liu, Dayiheng, Lin, Junyang, Huang, Fei, Zhou, Jingren
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910999827709952
author Zhang, Yanzhao
Li, Mingxin
Long, Dingkun
Zhang, Xin
Lin, Huan
Yang, Baosong
Xie, Pengjun
Yang, An
Liu, Dayiheng
Lin, Junyang
Huang, Fei
Zhou, Jingren
author_facet Zhang, Yanzhao
Li, Mingxin
Long, Dingkun
Zhang, Xin
Lin, Huan
Yang, Baosong
Xie, Pengjun
Yang, An
Liu, Dayiheng
Lin, Junyang
Huang, Fei
Zhou, Jingren
contents In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the Qwen3 LLMs serve not only as backbone models but also play a crucial role in synthesizing high-quality, rich, and diverse training data across multiple domains and languages, thus enhancing the training pipeline. The Qwen3 Embedding series offers a spectrum of model sizes (0.6B, 4B, 8B) for both embedding and reranking tasks, addressing diverse deployment scenarios where users can optimize for either efficiency or effectiveness. Empirical evaluations demonstrate that the Qwen3 Embedding series achieves state-of-the-art results across diverse benchmarks. Notably, it excels on the multilingual evaluation benchmark MTEB for text embedding, as well as in various retrieval tasks, including code retrieval, cross-lingual retrieval and multilingual retrieval. To facilitate reproducibility and promote community-driven research and development, the Qwen3 Embedding models are publicly available under the Apache 2.0 license.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05176
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Zhang, Yanzhao
Li, Mingxin
Long, Dingkun
Zhang, Xin
Lin, Huan
Yang, Baosong
Xie, Pengjun
Yang, An
Liu, Dayiheng
Lin, Junyang
Huang, Fei
Zhou, Jingren
Computation and Language
In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the Qwen3 LLMs serve not only as backbone models but also play a crucial role in synthesizing high-quality, rich, and diverse training data across multiple domains and languages, thus enhancing the training pipeline. The Qwen3 Embedding series offers a spectrum of model sizes (0.6B, 4B, 8B) for both embedding and reranking tasks, addressing diverse deployment scenarios where users can optimize for either efficiency or effectiveness. Empirical evaluations demonstrate that the Qwen3 Embedding series achieves state-of-the-art results across diverse benchmarks. Notably, it excels on the multilingual evaluation benchmark MTEB for text embedding, as well as in various retrieval tasks, including code retrieval, cross-lingual retrieval and multilingual retrieval. To facilitate reproducibility and promote community-driven research and development, the Qwen3 Embedding models are publicly available under the Apache 2.0 license.
title Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
topic Computation and Language
url https://arxiv.org/abs/2506.05176