MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiao, Zilin, Ma, Qi, Gu, Mengting, Chen, Chun-cheng Jason, Chen, Xintao, Ordonez, Vicente, Mohan, Vijai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917387143479296
author Xiao, Zilin
Ma, Qi
Gu, Mengting
Chen, Chun-cheng Jason
Chen, Xintao
Ordonez, Vicente
Mohan, Vijai
author_facet Xiao, Zilin
Ma, Qi
Gu, Mengting
Chen, Chun-cheng Jason
Chen, Xintao
Ordonez, Vicente
Mohan, Vijai
contents Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the expressiveness for fine-grained information, or produce too many vectors that are prohibitive for multi-vector retrieval. In this work, we introduce MetaEmbed, a new framework for multimodal retrieval that rethinks how multimodal embeddings are constructed and interacted with at scale. During training, a fixed number of learnable Meta Tokens are appended to the input sequence. At test-time, their last-layer contextualized representations serve as compact yet expressive multi-vector embeddings. Through the proposed Matryoshka Multi-Vector Retrieval training, MetaEmbed learns to organize information by granularity across multiple vectors. As a result, we enable test-time scaling in multimodal retrieval where users can balance retrieval quality against efficiency demands by selecting the number of tokens used for indexing and retrieval interactions. Extensive evaluations on the Massive Multimodal Embedding Benchmark (MMEB) and the Visual Document Retrieval Benchmark (ViDoRe) confirm that MetaEmbed achieves state-of-the-art retrieval performance while scaling robustly to models with 32B parameters. Code is available at https://github.com/facebookresearch/MetaEmbed.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
Xiao, Zilin
Ma, Qi
Gu, Mengting
Chen, Chun-cheng Jason
Chen, Xintao
Ordonez, Vicente
Mohan, Vijai
Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the expressiveness for fine-grained information, or produce too many vectors that are prohibitive for multi-vector retrieval. In this work, we introduce MetaEmbed, a new framework for multimodal retrieval that rethinks how multimodal embeddings are constructed and interacted with at scale. During training, a fixed number of learnable Meta Tokens are appended to the input sequence. At test-time, their last-layer contextualized representations serve as compact yet expressive multi-vector embeddings. Through the proposed Matryoshka Multi-Vector Retrieval training, MetaEmbed learns to organize information by granularity across multiple vectors. As a result, we enable test-time scaling in multimodal retrieval where users can balance retrieval quality against efficiency demands by selecting the number of tokens used for indexing and retrieval interactions. Extensive evaluations on the Massive Multimodal Embedding Benchmark (MMEB) and the Visual Document Retrieval Benchmark (ViDoRe) confirm that MetaEmbed achieves state-of-the-art retrieval performance while scaling robustly to models with 32B parameters. Code is available at https://github.com/facebookresearch/MetaEmbed.
title MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
topic Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18095