AMES: Approximate Multi-modal Enterprise Search via Late Interaction Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joseph, Tony, Pareja, Carlos, Pegna, David Lopes, Singh, Abhishek
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911514695303168
author Joseph, Tony
Pareja, Carlos
Pegna, David Lopes
Singh, Abhishek
author_facet Joseph, Tony
Pareja, Carlos
Pegna, David Lopes
Singh, Abhishek
contents We present AMES (Approximate Multimodal Enterprise Search), a unified multimodal late interaction retrieval architecture which is backend agnostic. AMES demonstrates that fine-grained multimodal late interaction retrieval can be deployed within a production grade enterprise search engine without architectural redesign. Text tokens, image patches, and video frames are embedded into a shared representation space using multi-vector encoders, enabling cross-modal retrieval without modality specific retrieval logic. AMES employs a two-stage pipeline: parallel token level ANN search with per document Top-M MaxSim approximation, followed by accelerator optimized Exact MaxSim re-ranking. Experiments on the ViDoRe V3 benchmark show that AMES achieves competitive ranking performance within a scalable, production ready Solr based system.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AMES: Approximate Multi-modal Enterprise Search via Late Interaction Retrieval
Joseph, Tony
Pareja, Carlos
Pegna, David Lopes
Singh, Abhishek
Information Retrieval
Machine Learning
We present AMES (Approximate Multimodal Enterprise Search), a unified multimodal late interaction retrieval architecture which is backend agnostic. AMES demonstrates that fine-grained multimodal late interaction retrieval can be deployed within a production grade enterprise search engine without architectural redesign. Text tokens, image patches, and video frames are embedded into a shared representation space using multi-vector encoders, enabling cross-modal retrieval without modality specific retrieval logic. AMES employs a two-stage pipeline: parallel token level ANN search with per document Top-M MaxSim approximation, followed by accelerator optimized Exact MaxSim re-ranking. Experiments on the ViDoRe V3 benchmark show that AMES achieves competitive ranking performance within a scalable, production ready Solr based system.
title AMES: Approximate Multi-modal Enterprise Search via Late Interaction Retrieval
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2603.13537