When Embedding Models Meet: Procrustes Bounds and Applications

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maystre, Lucas, Gonzalez, Alvaro Ortega, Park, Charles, Dolga, Rares, Berariu, Tudor, Zhao, Yu, Ciosek, Kamil
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911211165057024
author Maystre, Lucas
Gonzalez, Alvaro Ortega
Park, Charles
Dolga, Rares
Berariu, Tudor
Zhao, Yu
Ciosek, Kamil
author_facet Maystre, Lucas
Gonzalez, Alvaro Ortega
Park, Charles
Dolga, Rares
Berariu, Tudor
Zhao, Yu
Ciosek, Kamil
contents Embedding models trained separately on similar data often produce representations that encode stable information but are not directly interchangeable. This lack of interoperability raises challenges in several practical applications, such as model retraining, partial model upgrades, and multimodal search. Driven by these challenges, we study when two sets of embeddings can be aligned by an orthogonal transformation. We show that if pairwise dot products are approximately preserved, then there exists an isometry that closely aligns the two sets, and we provide a tight bound on the alignment error. This insight yields a simple alignment recipe, Procrustes post-processing, that makes two embedding models interoperable while preserving the geometry of each embedding space. Empirically, we demonstrate its effectiveness in three applications: maintaining compatibility across retrainings, combining different models for text retrieval, and improving mixed-modality search, where it achieves state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Embedding Models Meet: Procrustes Bounds and Applications
Maystre, Lucas
Gonzalez, Alvaro Ortega
Park, Charles
Dolga, Rares
Berariu, Tudor
Zhao, Yu
Ciosek, Kamil
Machine Learning
Embedding models trained separately on similar data often produce representations that encode stable information but are not directly interchangeable. This lack of interoperability raises challenges in several practical applications, such as model retraining, partial model upgrades, and multimodal search. Driven by these challenges, we study when two sets of embeddings can be aligned by an orthogonal transformation. We show that if pairwise dot products are approximately preserved, then there exists an isometry that closely aligns the two sets, and we provide a tight bound on the alignment error. This insight yields a simple alignment recipe, Procrustes post-processing, that makes two embedding models interoperable while preserving the geometry of each embedding space. Empirically, we demonstrate its effectiveness in three applications: maintaining compatibility across retrainings, combining different models for text retrieval, and improving mixed-modality search, where it achieves state-of-the-art performance.
title When Embedding Models Meet: Procrustes Bounds and Applications
topic Machine Learning
url https://arxiv.org/abs/2510.13406