Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsai, Yu-Che, Chen, Kuan-Yu, Li, Yuan-Chi, Chen, Yuan-Hao, Tsai, Ching-Yu, Lin, Shou-De
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914309830868992
author Tsai, Yu-Che
Chen, Kuan-Yu
Li, Yuan-Chi
Chen, Yuan-Hao
Tsai, Ching-Yu
Lin, Shou-De
author_facet Tsai, Yu-Che
Chen, Kuan-Yu
Li, Yuan-Chi
Chen, Yuan-Hao
Tsai, Ching-Yu
Lin, Shou-De
contents Existing large language model (LLM)-based embeddings typically adopt an encoder-only paradigm, treating LLMs as static feature extractors and overlooking their core generative strengths. We introduce GIRCSE (Generative Iterative Refinement for Contrastive Sentence Embeddings), a novel framework that leverages autoregressive generation to iteratively refine semantic representations. By producing sequences of soft tokens optimized under contrastive objective, GIRCSE captures latent concepts and implicit semantics that encoder-only methods often miss. To guide this process, we propose an Iterative Contrastive Refinement (ICR) objective that encourages each refinement step to yield better representations. Extensive experiments show that GIRCSE outperforms strong LLM-based embedding baselines on the MTEB benchmark and instruction-following tasks. Moreover, GIRCSE exhibits an emergent test-time scaling property: generating more tokens at inference steadily improves embedding quality. Our results establish generative iterative refinement as a new paradigm for representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
Tsai, Yu-Che
Chen, Kuan-Yu
Li, Yuan-Chi
Chen, Yuan-Hao
Tsai, Ching-Yu
Lin, Shou-De
Computation and Language
Artificial Intelligence
Existing large language model (LLM)-based embeddings typically adopt an encoder-only paradigm, treating LLMs as static feature extractors and overlooking their core generative strengths. We introduce GIRCSE (Generative Iterative Refinement for Contrastive Sentence Embeddings), a novel framework that leverages autoregressive generation to iteratively refine semantic representations. By producing sequences of soft tokens optimized under contrastive objective, GIRCSE captures latent concepts and implicit semantics that encoder-only methods often miss. To guide this process, we propose an Iterative Contrastive Refinement (ICR) objective that encourages each refinement step to yield better representations. Extensive experiments show that GIRCSE outperforms strong LLM-based embedding baselines on the MTEB benchmark and instruction-following tasks. Moreover, GIRCSE exhibits an emergent test-time scaling property: generating more tokens at inference steadily improves embedding quality. Our results establish generative iterative refinement as a new paradigm for representation learning.
title Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.24291