Little Giants: Synthesizing High-Quality Embedding Data at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Haonan, Wang, Liang, Yang, Nan, Zhu, Yutao, Zhao, Ziliang, Wei, Furu, Dou, Zhicheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917826602729472
author Chen, Haonan
Wang, Liang
Yang, Nan
Zhu, Yutao
Zhao, Ziliang
Wei, Furu
Dou, Zhicheng
author_facet Chen, Haonan
Wang, Liang
Yang, Nan
Zhu, Yutao
Zhao, Ziliang
Wei, Furu
Dou, Zhicheng
contents Synthetic data generation has become an increasingly popular way of training models without the need for large, manually labeled datasets. For tasks like text embedding, synthetic data offers diverse and scalable training examples, significantly reducing the cost of human annotation. However, most current approaches rely heavily on proprietary models like GPT-4, which are expensive and inefficient for generating large-scale embedding data. In this paper, we introduce SPEED, a framework that aligns open-source small models (8B) to efficiently generate large-scale synthetic embedding data. Through supervised fine-tuning, preference optimization, and self-improvement, SPEED enables small open-source models to produce high-quality data. Remarkably, SPEED uses only less than 1/10 of the GPT API calls, outperforming the state-of-the-art embedding model E5_mistral when both are trained solely on their synthetic data. Using this efficient generator, we conduct a comprehensive study on how various factors within the alignment pipeline impact data quality and reveal the scaling law for synthetic embedding data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18634
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Little Giants: Synthesizing High-Quality Embedding Data at Scale
Chen, Haonan
Wang, Liang
Yang, Nan
Zhu, Yutao
Zhao, Ziliang
Wei, Furu
Dou, Zhicheng
Computation and Language
Artificial Intelligence
Information Retrieval
Synthetic data generation has become an increasingly popular way of training models without the need for large, manually labeled datasets. For tasks like text embedding, synthetic data offers diverse and scalable training examples, significantly reducing the cost of human annotation. However, most current approaches rely heavily on proprietary models like GPT-4, which are expensive and inefficient for generating large-scale embedding data. In this paper, we introduce SPEED, a framework that aligns open-source small models (8B) to efficiently generate large-scale synthetic embedding data. Through supervised fine-tuning, preference optimization, and self-improvement, SPEED enables small open-source models to produce high-quality data. Remarkably, SPEED uses only less than 1/10 of the GPT API calls, outperforming the state-of-the-art embedding model E5_mistral when both are trained solely on their synthetic data. Using this efficient generator, we conduct a comprehensive study on how various factors within the alignment pipeline impact data quality and reveal the scaling law for synthetic embedding data.
title Little Giants: Synthesizing High-Quality Embedding Data at Scale
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2410.18634