Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Mengyao, Moreira, Gabriel, Ak, Ronay, Osmulski, Radek, Babakhin, Yauhen, Yu, Zhiding, Schifferer, Benedikt, Oldridge, Even
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916831805046784
author Xu, Mengyao
Moreira, Gabriel
Ak, Ronay
Osmulski, Radek
Babakhin, Yauhen
Yu, Zhiding
Schifferer, Benedikt
Oldridge, Even
author_facet Xu, Mengyao
Moreira, Gabriel
Ak, Ronay
Osmulski, Radek
Babakhin, Yauhen
Yu, Zhiding
Schifferer, Benedikt
Oldridge, Even
contents Motivated by the growing demand for retrieval systems that operate across modalities, we introduce llama-nemoretriever-colembed, a unified text-image retrieval model that delivers state-of-the-art performance across multiple benchmarks. We release two model variants, 1B and 3B. The 3B model achieves state of the art performance, scoring NDCG@5 91.0 on ViDoRe V1 and 63.5 on ViDoRe V2, placing first on both leaderboards as of June 27, 2025. Our approach leverages the NVIDIA Eagle2 Vision-Language model (VLM), modifies its architecture by replacing causal attention with bidirectional attention, and integrates a ColBERT-style late interaction mechanism to enable fine-grained multimodal retrieval in a shared embedding space. While this mechanism delivers superior retrieval accuracy, it introduces trade-offs in storage and efficiency. We provide a comprehensive analysis of these trade-offs. Additionally, we adopt a two-stage training strategy to enhance the model's retrieval capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
Xu, Mengyao
Moreira, Gabriel
Ak, Ronay
Osmulski, Radek
Babakhin, Yauhen
Yu, Zhiding
Schifferer, Benedikt
Oldridge, Even
Computer Vision and Pattern Recognition
Artificial Intelligence
Motivated by the growing demand for retrieval systems that operate across modalities, we introduce llama-nemoretriever-colembed, a unified text-image retrieval model that delivers state-of-the-art performance across multiple benchmarks. We release two model variants, 1B and 3B. The 3B model achieves state of the art performance, scoring NDCG@5 91.0 on ViDoRe V1 and 63.5 on ViDoRe V2, placing first on both leaderboards as of June 27, 2025. Our approach leverages the NVIDIA Eagle2 Vision-Language model (VLM), modifies its architecture by replacing causal attention with bidirectional attention, and integrates a ColBERT-style late interaction mechanism to enable fine-grained multimodal retrieval in a shared embedding space. While this mechanism delivers superior retrieval accuracy, it introduces trade-offs in storage and efficiency. We provide a comprehensive analysis of these trade-offs. Additionally, we adopt a two-stage training strategy to enhance the model's retrieval capabilities.
title Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.05513