LLM2Vec-Gen: Generative Embeddings from Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: BehnamGhader, Parishad, Adlakha, Vaibhav, Schmidt, Fabian David, Chapados, Nicolas, Mosbach, Marius, Reddy, Siva
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908933713559552
author BehnamGhader, Parishad
Adlakha, Vaibhav
Schmidt, Fabian David
Chapados, Nicolas
Mosbach, Marius
Reddy, Siva
author_facet BehnamGhader, Parishad
Adlakha, Vaibhav
Schmidt, Fabian David
Chapados, Nicolas
Mosbach, Marius
Reddy, Siva
contents Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead produces embeddings directly in the LLM's output space by learning to represent the model's potential response. Specifically, trainable special tokens are appended to the input and optimized to compress the LLM's own response into a fixed-length embedding, guided by an unsupervised embedding teacher and a reconstruction objective. Crucially, the LLM backbone remains frozen and training requires only unlabeled queries. LLM2Vec-Gen achieves state-of-the-art self-supervised performance on the Massive Text Embedding Benchmark (MTEB), improving by 8.8% over the unsupervised embedding teacher. Since the embeddings preserve the LLM's response-space semantics, they inherit capabilities such as safety alignment (up to 22.6% reduction in harmful content retrieval) and reasoning (up to 35.6% improvement on reasoning-intensive retrieval). Finally, the learned embeddings are also interpretable: they can be decoded back into text to reveal their semantic content.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM2Vec-Gen: Generative Embeddings from Large Language Models
BehnamGhader, Parishad
Adlakha, Vaibhav
Schmidt, Fabian David
Chapados, Nicolas
Mosbach, Marius
Reddy, Siva
Computation and Language
Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead produces embeddings directly in the LLM's output space by learning to represent the model's potential response. Specifically, trainable special tokens are appended to the input and optimized to compress the LLM's own response into a fixed-length embedding, guided by an unsupervised embedding teacher and a reconstruction objective. Crucially, the LLM backbone remains frozen and training requires only unlabeled queries. LLM2Vec-Gen achieves state-of-the-art self-supervised performance on the Massive Text Embedding Benchmark (MTEB), improving by 8.8% over the unsupervised embedding teacher. Since the embeddings preserve the LLM's response-space semantics, they inherit capabilities such as safety alignment (up to 22.6% reduction in harmful content retrieval) and reasoning (up to 35.6% improvement on reasoning-intensive retrieval). Finally, the learned embeddings are also interpretable: they can be decoded back into text to reveal their semantic content.
title LLM2Vec-Gen: Generative Embeddings from Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.10913