GRACE: Generative Representation Learning via Contrastive Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Jiashuo, Liu, Shixuan, Su, Zhaochen, Zhong, Xianrui, Jiang, Pengcheng, Jin, Bowen, Li, Peiran, Shi, Weijia, Han, Jiawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914076441968640
author Sun, Jiashuo
Liu, Shixuan
Su, Zhaochen
Zhong, Xianrui
Jiang, Pengcheng
Jin, Bowen
Li, Peiran
Shi, Weijia
Han, Jiawei
author_facet Sun, Jiashuo
Liu, Shixuan
Su, Zhaochen
Zhong, Xianrui
Jiang, Pengcheng
Jin, Bowen
Li, Peiran
Shi, Weijia
Han, Jiawei
contents Prevailing methods for training Large Language Models (LLMs) as text encoders rely on contrastive losses that treat the model as a black box function, discarding its generative and reasoning capabilities in favor of static embeddings. We introduce GRACE (Generative Representation Learning via Contrastive Policy Optimization), a novel framework that reimagines contrastive signals not as losses to be minimized, but as rewards that guide a generative policy. In GRACE, the LLM acts as a policy that produces explicit, human-interpretable rationales--structured natural language explanations of its semantic understanding. These rationales are then encoded into high-quality embeddings via mean pooling. Using policy gradient optimization, we train the model with a multi-component reward function that maximizes similarity between query positive pairs and minimizes similarity with negatives. This transforms the LLM from an opaque encoder into an interpretable agent whose reasoning process is transparent and inspectable. On MTEB benchmark, GRACE yields broad cross category gains: averaged over four backbones, the supervised setting improves overall score by 11.5% over base models, and the unsupervised variant adds 6.9%, while preserving general capabilities. This work treats contrastive objectives as rewards over rationales, unifying representation learning with generation to produce stronger embeddings and transparent rationales. The model, data and code are available at https://github.com/GasolSun36/GRACE.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRACE: Generative Representation Learning via Contrastive Policy Optimization
Sun, Jiashuo
Liu, Shixuan
Su, Zhaochen
Zhong, Xianrui
Jiang, Pengcheng
Jin, Bowen
Li, Peiran
Shi, Weijia
Han, Jiawei
Computation and Language
Artificial Intelligence
Information Retrieval
Prevailing methods for training Large Language Models (LLMs) as text encoders rely on contrastive losses that treat the model as a black box function, discarding its generative and reasoning capabilities in favor of static embeddings. We introduce GRACE (Generative Representation Learning via Contrastive Policy Optimization), a novel framework that reimagines contrastive signals not as losses to be minimized, but as rewards that guide a generative policy. In GRACE, the LLM acts as a policy that produces explicit, human-interpretable rationales--structured natural language explanations of its semantic understanding. These rationales are then encoded into high-quality embeddings via mean pooling. Using policy gradient optimization, we train the model with a multi-component reward function that maximizes similarity between query positive pairs and minimizes similarity with negatives. This transforms the LLM from an opaque encoder into an interpretable agent whose reasoning process is transparent and inspectable. On MTEB benchmark, GRACE yields broad cross category gains: averaged over four backbones, the supervised setting improves overall score by 11.5% over base models, and the unsupervised variant adds 6.9%, while preserving general capabilities. This work treats contrastive objectives as rewards over rationales, unifying representation learning with generation to produce stronger embeddings and transparent rationales. The model, data and code are available at https://github.com/GasolSun36/GRACE.
title GRACE: Generative Representation Learning via Contrastive Policy Optimization
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2510.04506