Understanding the Influence of Synthetic Data for Text Embedders

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Springer, Jacob Mitchell, Adlakha, Vaibhav, Reddy, Siva, Raghunathan, Aditi, Mosbach, Marius
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916939151966208
author Springer, Jacob Mitchell
Adlakha, Vaibhav
Reddy, Siva
Raghunathan, Aditi
Mosbach, Marius
author_facet Springer, Jacob Mitchell
Adlakha, Vaibhav
Reddy, Siva
Raghunathan, Aditi
Mosbach, Marius
contents Recent progress in developing general purpose text embedders has been driven by training on ever-growing corpora of synthetic LLM-generated data. Nonetheless, no publicly available synthetic dataset exists, posing a barrier to studying its role for generalization. To address this issue, we first reproduce and publicly release the synthetic data proposed by Wang et al. (Mistral-E5). Our synthetic data is high quality and leads to consistent improvements in performance. Next, we critically examine where exactly synthetic data improves model generalization. Our analysis reveals that benefits from synthetic data are sparse and highly localized to individual datasets. Moreover, we observe trade-offs between the performance on different categories and data that benefits one task, degrades performance on another. Our findings highlight the limitations of current synthetic data approaches for building general-purpose embedders and challenge the notion that training on synthetic data leads to more robust embedding models across tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the Influence of Synthetic Data for Text Embedders
Springer, Jacob Mitchell
Adlakha, Vaibhav
Reddy, Siva
Raghunathan, Aditi
Mosbach, Marius
Computation and Language
Recent progress in developing general purpose text embedders has been driven by training on ever-growing corpora of synthetic LLM-generated data. Nonetheless, no publicly available synthetic dataset exists, posing a barrier to studying its role for generalization. To address this issue, we first reproduce and publicly release the synthetic data proposed by Wang et al. (Mistral-E5). Our synthetic data is high quality and leads to consistent improvements in performance. Next, we critically examine where exactly synthetic data improves model generalization. Our analysis reveals that benefits from synthetic data are sparse and highly localized to individual datasets. Moreover, we observe trade-offs between the performance on different categories and data that benefits one task, degrades performance on another. Our findings highlight the limitations of current synthetic data approaches for building general-purpose embedders and challenge the notion that training on synthetic data leads to more robust embedding models across tasks.
title Understanding the Influence of Synthetic Data for Text Embedders
topic Computation and Language
url https://arxiv.org/abs/2509.06184