Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tamber, Manveer Singh, Xian, Jasper, Lin, Jimmy
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918007754719232
author Tamber, Manveer Singh
Xian, Jasper
Lin, Jimmy
author_facet Tamber, Manveer Singh
Xian, Jasper
Lin, Jimmy
contents Embedding models that generate dense vector representations of text are widely used and hold significant commercial value. Companies such as OpenAI and Cohere offer proprietary embedding models via paid APIs, but despite being "hidden" behind APIs, these models are not protected from theft. We present, to our knowledge, the first effort to "steal" these models for retrieval by training thief models on text-embedding pairs obtained from the APIs. Our experiments demonstrate that it is possible to replicate the retrieval effectiveness of commercial embedding models with a cost of under $300. Notably, our methods allow for distilling from multiple teachers into a single robust student model, and for distilling into presumably smaller models with fewer dimension vectors, yet competitive retrieval effectiveness. Our findings raise important considerations for deploying commercial embedding models and suggest measures to mitigate the risk of model theft.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09355
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
Tamber, Manveer Singh
Xian, Jasper
Lin, Jimmy
Information Retrieval
Embedding models that generate dense vector representations of text are widely used and hold significant commercial value. Companies such as OpenAI and Cohere offer proprietary embedding models via paid APIs, but despite being "hidden" behind APIs, these models are not protected from theft. We present, to our knowledge, the first effort to "steal" these models for retrieval by training thief models on text-embedding pairs obtained from the APIs. Our experiments demonstrate that it is possible to replicate the retrieval effectiveness of commercial embedding models with a cost of under $300. Notably, our methods allow for distilling from multiple teachers into a single robust student model, and for distilling into presumably smaller models with fewer dimension vectors, yet competitive retrieval effectiveness. Our findings raise important considerations for deploying commercial embedding models and suggest measures to mitigate the risk of model theft.
title Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
topic Information Retrieval
url https://arxiv.org/abs/2406.09355