Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Lang, Wu, Qiyu, Miao, Zhongtao, Yamasaki, Toshihiko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912250362593280
author Huang, Lang
Wu, Qiyu
Miao, Zhongtao
Yamasaki, Toshihiko
author_facet Huang, Lang
Wu, Qiyu
Miao, Zhongtao
Yamasaki, Toshihiko
contents Information retrieval is indispensable for today's Internet applications, yet traditional semantic matching techniques often fall short in capturing the fine-grained cross-modal interactions required for complex queries. Although late-fusion two-tower architectures attempt to bridge this gap by independently encoding visual and textual data before merging them at a high level, they frequently overlook the subtle interplay essential for comprehensive understanding. In this work, we rigorously assess these limitations and introduce a unified retrieval framework that fuses visual and textual cues from the ground up, enabling early cross-modal interactions for enhancing context interpretation. Through a two-stage training process--comprising post-training adaptation followed by instruction tuning--we adapt MLLMs as retrievers using a simple one-tower architecture. Our approach outperforms conventional methods across diverse retrieval scenarios, particularly when processing complex multi-modal inputs. Notably, the joint fusion encoder yields greater improvements on tasks that require modality fusion compared to those that do not, underscoring the transformative potential of early integration strategies and pointing toward a promising direction for contextually aware and effective information retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
Huang, Lang
Wu, Qiyu
Miao, Zhongtao
Yamasaki, Toshihiko
Computer Vision and Pattern Recognition
Information retrieval is indispensable for today's Internet applications, yet traditional semantic matching techniques often fall short in capturing the fine-grained cross-modal interactions required for complex queries. Although late-fusion two-tower architectures attempt to bridge this gap by independently encoding visual and textual data before merging them at a high level, they frequently overlook the subtle interplay essential for comprehensive understanding. In this work, we rigorously assess these limitations and introduce a unified retrieval framework that fuses visual and textual cues from the ground up, enabling early cross-modal interactions for enhancing context interpretation. Through a two-stage training process--comprising post-training adaptation followed by instruction tuning--we adapt MLLMs as retrievers using a simple one-tower architecture. Our approach outperforms conventional methods across diverse retrieval scenarios, particularly when processing complex multi-modal inputs. Notably, the joint fusion encoder yields greater improvements on tasks that require modality fusion compared to those that do not, underscoring the transformative potential of early integration strategies and pointing toward a promising direction for contextually aware and effective information retrieval.
title Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.20008