Recurrence Meets Transformers for Universal Multimodal Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caffagni, Davide, Sarto, Sara, Cornia, Marcella, Baraldi, Lorenzo, Cucchiara, Rita
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916945948835840
author Caffagni, Davide
Sarto, Sara
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
author_facet Caffagni, Davide
Sarto, Sara
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
contents With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks have emerged. Existing methods predominantly rely on task-specific fine-tuning of vision-language models and are limited to single-modality queries or documents. In this paper, we propose ReT-2, a unified retrieval model that supports multimodal queries, composed of both images and text, and searches across multimodal document collections where text and images coexist. ReT-2 leverages multi-layer representations and a recurrent Transformer architecture with LSTM-inspired gating mechanisms to dynamically integrate information across layers and modalities, capturing fine-grained visual and textual details. We evaluate ReT-2 on the challenging M2KR and M-BEIR benchmarks across different retrieval configurations. Results demonstrate that ReT-2 consistently achieves state-of-the-art performance across diverse settings, while offering faster inference and reduced memory usage compared to prior approaches. When integrated into retrieval-augmented generation pipelines, ReT-2 also improves downstream performance on Encyclopedic-VQA and InfoSeek datasets. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT-2
format Preprint
id arxiv_https___arxiv_org_abs_2509_08897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recurrence Meets Transformers for Universal Multimodal Retrieval
Caffagni, Davide
Sarto, Sara
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks have emerged. Existing methods predominantly rely on task-specific fine-tuning of vision-language models and are limited to single-modality queries or documents. In this paper, we propose ReT-2, a unified retrieval model that supports multimodal queries, composed of both images and text, and searches across multimodal document collections where text and images coexist. ReT-2 leverages multi-layer representations and a recurrent Transformer architecture with LSTM-inspired gating mechanisms to dynamically integrate information across layers and modalities, capturing fine-grained visual and textual details. We evaluate ReT-2 on the challenging M2KR and M-BEIR benchmarks across different retrieval configurations. Results demonstrate that ReT-2 consistently achieves state-of-the-art performance across diverse settings, while offering faster inference and reduced memory usage compared to prior approaches. When integrated into retrieval-augmented generation pipelines, ReT-2 also improves downstream performance on Encyclopedic-VQA and InfoSeek datasets. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT-2
title Recurrence Meets Transformers for Universal Multimodal Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2509.08897