Text2Token: Unsupervised Text Representation Learning with Token Target Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: An, Ruize, Zhang, Richong, Nie, Zhijie, Wu, Zhanyu, Zhang, Yanzhao, Long, Dingkun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909837687783424
author An, Ruize
Zhang, Richong
Nie, Zhijie
Wu, Zhanyu
Zhang, Yanzhao
Long, Dingkun
author_facet An, Ruize
Zhang, Richong
Nie, Zhijie
Wu, Zhanyu
Zhang, Yanzhao
Long, Dingkun
contents Unsupervised text representation learning (TRL) is a fundamental task in natural language processing, which is beneficial for improving search and recommendations with the web's unlabeled texts. A recent empirical study finds that the high-quality representation aligns with the key token of the input text, uncovering the potential connection between representation space and vocabulary space. Inspired by the findings, we revisit the generative tasks and develop an unsupervised generative framework for TRL, Text2Token. The framework is based on the token target prediction task, utilizing carefully constructed target token distribution as supervisory signals. To construct the high-quality target token distribution, we analyze the token-alignment properties with advanced embedders and identify two essential categories of key tokens: (1) the meaningful tokens in the text and (2) semantically derived tokens beyond the text. Based on these insights, we propose two methods -- data-driven and model-derived -- to construct synthetic token targets from data or the LLM backbone. Experiments on the MTEB v2 benchmark demonstrate that Text2Token achieves performance competitive with the state-of-the-art embedder with unsupervised contrastive learning, LLM2Vec. Our analysis further shows that vocabulary and representation spaces optimize together and toward the optimum solution during training, providing new ideas and insights for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
An, Ruize
Zhang, Richong
Nie, Zhijie
Wu, Zhanyu
Zhang, Yanzhao
Long, Dingkun
Computation and Language
Information Retrieval
Unsupervised text representation learning (TRL) is a fundamental task in natural language processing, which is beneficial for improving search and recommendations with the web's unlabeled texts. A recent empirical study finds that the high-quality representation aligns with the key token of the input text, uncovering the potential connection between representation space and vocabulary space. Inspired by the findings, we revisit the generative tasks and develop an unsupervised generative framework for TRL, Text2Token. The framework is based on the token target prediction task, utilizing carefully constructed target token distribution as supervisory signals. To construct the high-quality target token distribution, we analyze the token-alignment properties with advanced embedders and identify two essential categories of key tokens: (1) the meaningful tokens in the text and (2) semantically derived tokens beyond the text. Based on these insights, we propose two methods -- data-driven and model-derived -- to construct synthetic token targets from data or the LLM backbone. Experiments on the MTEB v2 benchmark demonstrate that Text2Token achieves performance competitive with the state-of-the-art embedder with unsupervised contrastive learning, LLM2Vec. Our analysis further shows that vocabulary and representation spaces optimize together and toward the optimum solution during training, providing new ideas and insights for future work.
title Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2510.10224