CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huo, Jiahao, Huang, Yu, Yan, Yibo, Pan, Ye, Zheng, Kening, Huang, Wei-Chieh, Cao, Yi, Ou, Mingdong, Yu, Philip S., Hu, Xuming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910134068838400
author Huo, Jiahao
Huang, Yu
Yan, Yibo
Pan, Ye
Zheng, Kening
Huang, Wei-Chieh
Cao, Yi
Ou, Mingdong
Yu, Philip S.
Hu, Xuming
author_facet Huo, Jiahao
Huang, Yu
Yan, Yibo
Pan, Ye
Zheng, Kening
Huang, Wei-Chieh
Cao, Yi
Ou, Mingdong
Yu, Philip S.
Hu, Xuming
contents Although Multimodal Large Language Models (MLLMs) have shown remarkable potential in Visual Document Retrieval (VDR) through generating high-quality multi-vector embeddings, the substantial storage overhead caused by representing a page with thousands of visual tokens limits their practicality in real-world applications. To address this challenge, we propose an auto-regressive generation approach, CausalEmbed, for constructing multi-vector embeddings. By incorporating iterative margin loss during contrastive training, CausalEmbed encourages the embedding models to learn compact and well-structured representations. Our method enables efficient VDR tasks using only dozens of visual tokens, achieving a 30-155x reduction in token count while maintaining highly competitive performance across various backbones and benchmarks. Theoretical analysis and empirical results demonstrate the unique advantages of auto-regressive embedding generation in terms of training efficiency and scalability at test time. As a result, CausalEmbed introduces a flexible test-time scaling strategy for multi-vector VDR representations and sheds light on the generative paradigm within multimodal document retrieval. Our code is available at https://github.com/Z1zs/Causal-Embed.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21262
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
Huo, Jiahao
Huang, Yu
Yan, Yibo
Pan, Ye
Zheng, Kening
Huang, Wei-Chieh
Cao, Yi
Ou, Mingdong
Yu, Philip S.
Hu, Xuming
Computation and Language
Although Multimodal Large Language Models (MLLMs) have shown remarkable potential in Visual Document Retrieval (VDR) through generating high-quality multi-vector embeddings, the substantial storage overhead caused by representing a page with thousands of visual tokens limits their practicality in real-world applications. To address this challenge, we propose an auto-regressive generation approach, CausalEmbed, for constructing multi-vector embeddings. By incorporating iterative margin loss during contrastive training, CausalEmbed encourages the embedding models to learn compact and well-structured representations. Our method enables efficient VDR tasks using only dozens of visual tokens, achieving a 30-155x reduction in token count while maintaining highly competitive performance across various backbones and benchmarks. Theoretical analysis and empirical results demonstrate the unique advantages of auto-regressive embedding generation in terms of training efficiency and scalability at test time. As a result, CausalEmbed introduces a flexible test-time scaling strategy for multi-vector VDR representations and sheds light on the generative paradigm within multimodal document retrieval. Our code is available at https://github.com/Z1zs/Causal-Embed.
title CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
topic Computation and Language
url https://arxiv.org/abs/2601.21262