Do Reasoning Models Enhance Embedding Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chan, Wun Yu, Chen, Shaojin, Jing, Huihao, Lau, Kwun Hang, Li, Elton Chun-Chai, Wang, Zihao, Li, Haoran, Song, Yangqiu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914289350082560
author Chan, Wun Yu
Chen, Shaojin
Jing, Huihao
Lau, Kwun Hang
Li, Elton Chun-Chai
Wang, Zihao
Li, Haoran
Song, Yangqiu
author_facet Chan, Wun Yu
Chen, Shaojin
Jing, Huihao
Lau, Kwun Hang
Li, Elton Chun-Chai
Wang, Zihao
Li, Haoran
Song, Yangqiu
contents State-of-the-art embedding models are increasingly derived from decoder-only Large Language Model (LLM) backbones adapted via contrastive learning. Given the emergence of reasoning models trained via Reinforcement Learning with Verifiable Rewards (RLVR), a natural question arises: do enhanced reasoning translate to superior semantic representations when these models serve as embedding initializations? Contrary to expectation, our evaluation on MTEB and BRIGHT reveals a **null effect**: embedding models initialized from RLVR-tuned backbones yield no consistent performance advantage over their base counterparts when subjected to identical training recipes. To unpack this paradox, we introduce **H**ierarchical **R**epresentation **S**imilarity **A**nalysis (HRSA), a framework that decomposes similarity across representation, geometry, and function levels. HRSA reveals that while RLVR induces irreversible latent manifold's local geometry reorganization and reversible coordinate basis drift, it preserves the global manifold geometry and linear readout. Consequently, subsequent contrastive learning drives strong alignment between base- and reasoning-initialized models, a phenomenon we term **Manifold Realignment**. Empirically, our findings suggest that unlike Supervised Fine-Tuning (SFT), RLVR optimizes trajectories within an existing semantic landscape rather than fundamentally restructuring the landscape itself.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21192
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Reasoning Models Enhance Embedding Models?
Chan, Wun Yu
Chen, Shaojin
Jing, Huihao
Lau, Kwun Hang
Li, Elton Chun-Chai
Wang, Zihao
Li, Haoran
Song, Yangqiu
Artificial Intelligence
Computation and Language
68T07 (Primary) 68T50 (Secondary)
I.2.7; I.2.6
State-of-the-art embedding models are increasingly derived from decoder-only Large Language Model (LLM) backbones adapted via contrastive learning. Given the emergence of reasoning models trained via Reinforcement Learning with Verifiable Rewards (RLVR), a natural question arises: do enhanced reasoning translate to superior semantic representations when these models serve as embedding initializations? Contrary to expectation, our evaluation on MTEB and BRIGHT reveals a **null effect**: embedding models initialized from RLVR-tuned backbones yield no consistent performance advantage over their base counterparts when subjected to identical training recipes. To unpack this paradox, we introduce **H**ierarchical **R**epresentation **S**imilarity **A**nalysis (HRSA), a framework that decomposes similarity across representation, geometry, and function levels. HRSA reveals that while RLVR induces irreversible latent manifold's local geometry reorganization and reversible coordinate basis drift, it preserves the global manifold geometry and linear readout. Consequently, subsequent contrastive learning drives strong alignment between base- and reasoning-initialized models, a phenomenon we term **Manifold Realignment**. Empirically, our findings suggest that unlike Supervised Fine-Tuning (SFT), RLVR optimizes trajectories within an existing semantic landscape rather than fundamentally restructuring the landscape itself.
title Do Reasoning Models Enhance Embedding Models?
topic Artificial Intelligence
Computation and Language
68T07 (Primary) 68T50 (Secondary)
I.2.7; I.2.6
url https://arxiv.org/abs/2601.21192