TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Xiangzhao, Wang, Shijie, Yang, Tianyu, Wang, Tianyue, Guo, Haiyun, Wang, Jinqiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912941082673152
author Hao, Xiangzhao
Wang, Shijie
Yang, Tianyu
Wang, Tianyue
Guo, Haiyun
Wang, Jinqiao
author_facet Hao, Xiangzhao
Wang, Shijie
Yang, Tianyu
Wang, Tianyue
Guo, Haiyun
Wang, Jinqiao
contents Universal Multimodal Retrieval requires unified embedding models capable of interpreting diverse user intents, ranging from simple keywords to complex compositional instructions. While Multimodal Large Language Models (MLLMs) possess strong reasoning capabilities, prevailing adaptations confine them to static encoders, underutilizing their generative potential. This encoder-only paradigm struggles with complex intents that demand logical deduction rather than superficial pattern matching. To address this, we introduce TRACE (Task-adaptive Reasoning And Compressing Embeddings). TRACE unifies generative reasoning with discriminative representation learning. It first generates a structured Chain-of-Thought (CoT) to explicitly reason about the query, and subsequently compresses this reasoning trace into a compact embedding via a dedicated token. To train this framework, we construct M-BEIR-CoT, a large-scale dataset featuring a difficulty-aware routing strategy. Experiments on the M-BEIR benchmark establish TRACE as the new state-of-the-art. Crucially, TRACE demonstrates a learned implicit routing behavior. It autonomously activates reasoning for complex queries while bypassing it for simpler ones, achieving an optimal balance between retrieval accuracy and inference throughput. Furthermore, by internalizing the deductive process, TRACE exhibits remarkable zero-shot transferability to unseen domains and novel constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02929
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
Hao, Xiangzhao
Wang, Shijie
Yang, Tianyu
Wang, Tianyue
Guo, Haiyun
Wang, Jinqiao
Computer Vision and Pattern Recognition
Universal Multimodal Retrieval requires unified embedding models capable of interpreting diverse user intents, ranging from simple keywords to complex compositional instructions. While Multimodal Large Language Models (MLLMs) possess strong reasoning capabilities, prevailing adaptations confine them to static encoders, underutilizing their generative potential. This encoder-only paradigm struggles with complex intents that demand logical deduction rather than superficial pattern matching. To address this, we introduce TRACE (Task-adaptive Reasoning And Compressing Embeddings). TRACE unifies generative reasoning with discriminative representation learning. It first generates a structured Chain-of-Thought (CoT) to explicitly reason about the query, and subsequently compresses this reasoning trace into a compact embedding via a dedicated token. To train this framework, we construct M-BEIR-CoT, a large-scale dataset featuring a difficulty-aware routing strategy. Experiments on the M-BEIR benchmark establish TRACE as the new state-of-the-art. Crucially, TRACE demonstrates a learned implicit routing behavior. It autonomously activates reasoning for complex queries while bypassing it for simpler ones, achieving an optimal balance between retrieval accuracy and inference throughput. Furthermore, by internalizing the deductive process, TRACE exhibits remarkable zero-shot transferability to unseen domains and novel constraints.
title TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.02929