GEM: Empowering LLM for both Embedding Generation and Language Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Caojin, Zhang, Qiang, Li, Ke, Nuthalapati, Sai Vidyaranya, Zhang, Benyu, Liu, Jason, Li, Serena, Zhang, Lizhu, Fan, Xiangjun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916780381831168
author Zhang, Caojin
Zhang, Qiang
Li, Ke
Nuthalapati, Sai Vidyaranya
Zhang, Benyu
Liu, Jason
Li, Serena
Zhang, Lizhu
Fan, Xiangjun
author_facet Zhang, Caojin
Zhang, Qiang
Li, Ke
Nuthalapati, Sai Vidyaranya
Zhang, Benyu
Liu, Jason
Li, Serena
Zhang, Lizhu
Fan, Xiangjun
contents Large decoder-only language models (LLMs) have achieved remarkable success in generation and reasoning tasks, where they generate text responses given instructions. However, many applications, e.g., retrieval augmented generation (RAG), still rely on separate embedding models to generate text embeddings, which can complicate the system and introduce discrepancies in understanding of the query between the embedding model and LLMs. To address this limitation, we propose a simple self-supervised approach, Generative Embedding large language Model (GEM), that enables any large decoder-only LLM to generate high-quality text embeddings while maintaining its original text generation and reasoning capabilities. Our method inserts new special token(s) into a text body, and generates summarization embedding of the text by manipulating the attention mask. This method could be easily integrated into post-training or fine tuning stages of any existing LLMs. We demonstrate the effectiveness of our approach by applying it to two popular LLM families, ranging from 1B to 8B parameters, and evaluating the transformed models on both text embedding benchmarks (MTEB) and NLP benchmarks (MMLU). The results show that our proposed method significantly improves the original LLMs on MTEB while having a minimal impact on MMLU. Our strong results indicate that our approach can empower LLMs with state-of-the-art text embedding capabilities while maintaining their original NLP performance
format Preprint
id arxiv_https___arxiv_org_abs_2506_04344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GEM: Empowering LLM for both Embedding Generation and Language Understanding
Zhang, Caojin
Zhang, Qiang
Li, Ke
Nuthalapati, Sai Vidyaranya
Zhang, Benyu
Liu, Jason
Li, Serena
Zhang, Lizhu
Fan, Xiangjun
Computation and Language
Machine Learning
Large decoder-only language models (LLMs) have achieved remarkable success in generation and reasoning tasks, where they generate text responses given instructions. However, many applications, e.g., retrieval augmented generation (RAG), still rely on separate embedding models to generate text embeddings, which can complicate the system and introduce discrepancies in understanding of the query between the embedding model and LLMs. To address this limitation, we propose a simple self-supervised approach, Generative Embedding large language Model (GEM), that enables any large decoder-only LLM to generate high-quality text embeddings while maintaining its original text generation and reasoning capabilities. Our method inserts new special token(s) into a text body, and generates summarization embedding of the text by manipulating the attention mask. This method could be easily integrated into post-training or fine tuning stages of any existing LLMs. We demonstrate the effectiveness of our approach by applying it to two popular LLM families, ranging from 1B to 8B parameters, and evaluating the transformed models on both text embedding benchmarks (MTEB) and NLP benchmarks (MMLU). The results show that our proposed method significantly improves the original LLMs on MTEB while having a minimal impact on MMLU. Our strong results indicate that our approach can empower LLMs with state-of-the-art text embedding capabilities while maintaining their original NLP performance
title GEM: Empowering LLM for both Embedding Generation and Language Understanding
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.04344