The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Junchen, Ge, Xuri, Xin, Xin, Karatzoglou, Alexandros, Arapakis, Ioannis, Wang, Xi, Liu, Qijiong, Li, Qian, Jose, Joemon M.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911720038989824
author Fu, Junchen
Ge, Xuri
Xin, Xin
Karatzoglou, Alexandros
Arapakis, Ioannis
Wang, Xi
Liu, Qijiong
Li, Qian
Jose, Joemon M.
author_facet Fu, Junchen
Ge, Xuri
Xin, Xin
Karatzoglou, Alexandros
Arapakis, Ioannis
Wang, Xi
Liu, Qijiong
Li, Qian
Jose, Joemon M.
contents Multimodal representation learning has attracted increasing attention in AI, driven by the strong performance of large, pretrained multimodal foundation models such as Qwen, LLaVA, and CLIP. These models deliver impressive performance on a range of multimodal information retrieval (MIR) tasks, including web search, cross-modal retrieval, and recommender systems. Yet their massive parameter counts create major efficiency bottlenecks when adapting their representations for IR tasks during training, deployment, and inference. These limitations hinder the practical use of foundation models for representation learning in information retrieval. To address these issues, we propose organizing the EReL@MIR workshop at MM 2026, bringing together researchers from academia and industry to discuss emerging solutions, open challenges, and new efficiency metrics and benchmarks for multimodal IR representation learning in the foundation-model era. The workshop's official website is available at https://erel-mir.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26941
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
Fu, Junchen
Ge, Xuri
Xin, Xin
Karatzoglou, Alexandros
Arapakis, Ioannis
Wang, Xi
Liu, Qijiong
Li, Qian
Jose, Joemon M.
Information Retrieval
Multimedia
Multimodal representation learning has attracted increasing attention in AI, driven by the strong performance of large, pretrained multimodal foundation models such as Qwen, LLaVA, and CLIP. These models deliver impressive performance on a range of multimodal information retrieval (MIR) tasks, including web search, cross-modal retrieval, and recommender systems. Yet their massive parameter counts create major efficiency bottlenecks when adapting their representations for IR tasks during training, deployment, and inference. These limitations hinder the practical use of foundation models for representation learning in information retrieval. To address these issues, we propose organizing the EReL@MIR workshop at MM 2026, bringing together researchers from academia and industry to discuss emerging solutions, open challenges, and new efficiency metrics and benchmarks for multimodal IR representation learning in the foundation-model era. The workshop's official website is available at https://erel-mir.github.io/.
title The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
topic Information Retrieval
Multimedia
url https://arxiv.org/abs/2605.26941