Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Moreira, Gabriel de Souza P., Ak, Ronay, Xu, Mengyao, Holworthy, Oliver, Schifferer, Benedikt, Yu, Zhiding, Babakhin, Yauhen, Osmulski, Radek, Cai, Jiarui, Chesler, Ryan, Liu, Bo, Oldridge, Even
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917376631504896
author Moreira, Gabriel de Souza P.
Ak, Ronay
Xu, Mengyao
Holworthy, Oliver
Schifferer, Benedikt
Yu, Zhiding
Babakhin, Yauhen
Osmulski, Radek
Cai, Jiarui
Chesler, Ryan
Liu, Bo
Oldridge, Even
author_facet Moreira, Gabriel de Souza P.
Ak, Ronay
Xu, Mengyao
Holworthy, Oliver
Schifferer, Benedikt
Yu, Zhiding
Babakhin, Yauhen
Osmulski, Radek
Cai, Jiarui
Chesler, Ryan
Liu, Bo
Oldridge, Even
contents Retrieval-Augmented Generation (RAG) systems have been popular for generative applications, powering language models by injecting external knowledge. Companies have been trying to leverage their large catalog of documents (e.g. PDFs, presentation slides) in such RAG pipelines, whose first step is the retrieval component. Dense retrieval has been a popular approach, where embedding models are used to generate a dense representation of the user query that is closer to relevant content embeddings. More recently, VLM-based embedding models have become popular for visual document retrieval, as they preserve visual information and simplify the indexing pipeline compared to OCR text extraction. Motivated by the growing demand for visual document retrieval, we introduce Nemotron ColEmbed V2, a family of models that achieve state-of-the-art performance on the ViDoRe benchmarks. We release three variants - with 3B, 4B, and 8B parameters - based on pre-trained VLMs: NVIDIA Eagle 2 with Llama 3.2 3B backbone, Qwen3-VL-4B-Instruct and Qwen3-VL-8B-Instruct, respectively. The 8B model ranks first on the ViDoRe V3 leaderboard as of February 03, 2026, achieving an average NDCG@10 of 63.42. We describe the main techniques used across data processing, training, and post-training - such as cluster-based sampling, hard-negative mining, bidirectional attention, late interaction, and model merging - that helped us build our top-performing models. We also discuss compute and storage engineering challenges posed by the late interaction mechanism and present experiments on how to balance accuracy and storage with lower dimension embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03992
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
Moreira, Gabriel de Souza P.
Ak, Ronay
Xu, Mengyao
Holworthy, Oliver
Schifferer, Benedikt
Yu, Zhiding
Babakhin, Yauhen
Osmulski, Radek
Cai, Jiarui
Chesler, Ryan
Liu, Bo
Oldridge, Even
Information Retrieval
Retrieval-Augmented Generation (RAG) systems have been popular for generative applications, powering language models by injecting external knowledge. Companies have been trying to leverage their large catalog of documents (e.g. PDFs, presentation slides) in such RAG pipelines, whose first step is the retrieval component. Dense retrieval has been a popular approach, where embedding models are used to generate a dense representation of the user query that is closer to relevant content embeddings. More recently, VLM-based embedding models have become popular for visual document retrieval, as they preserve visual information and simplify the indexing pipeline compared to OCR text extraction. Motivated by the growing demand for visual document retrieval, we introduce Nemotron ColEmbed V2, a family of models that achieve state-of-the-art performance on the ViDoRe benchmarks. We release three variants - with 3B, 4B, and 8B parameters - based on pre-trained VLMs: NVIDIA Eagle 2 with Llama 3.2 3B backbone, Qwen3-VL-4B-Instruct and Qwen3-VL-8B-Instruct, respectively. The 8B model ranks first on the ViDoRe V3 leaderboard as of February 03, 2026, achieving an average NDCG@10 of 63.42. We describe the main techniques used across data processing, training, and post-training - such as cluster-based sampling, hard-negative mining, bidirectional attention, late interaction, and model merging - that helped us build our top-performing models. We also discuss compute and storage engineering challenges posed by the late interaction mechanism and present experiments on how to balance accuracy and storage with lower dimension embeddings.
title Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2602.03992