The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Junchen, Ge, Xuri, Xin, Xin, Yu, Haitao, Feng, Yue, Karatzoglou, Alexandros, Arapakis, Ioannis, Jose, Joemon M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917992565047296
author Fu, Junchen
Ge, Xuri
Xin, Xin
Yu, Haitao
Feng, Yue
Karatzoglou, Alexandros
Arapakis, Ioannis
Jose, Joemon M.
author_facet Fu, Junchen
Ge, Xuri
Xin, Xin
Yu, Haitao
Feng, Yue
Karatzoglou, Alexandros
Arapakis, Ioannis
Jose, Joemon M.
contents Multimodal representation learning has garnered significant attention in the AI community, largely due to the success of large pre-trained multimodal foundation models like LLaMA, GPT, Mistral, and CLIP. These models have achieved remarkable performance across various tasks of multimodal information retrieval (MIR), including web search, cross-modal retrieval, and recommender systems, etc. However, due to their enormous parameter sizes, significant efficiency challenges emerge across training, deployment, and inference stages when adapting these models' representation for IR tasks. These challenges present substantial obstacles to the practical adaptation of foundation models for representation learning in information retrieval tasks. To address these pressing issues, we propose organizing the first EReL@MIR workshop at the Web Conference 2025, inviting participants to explore novel solutions, emerging problems, challenges, efficiency evaluation metrics and benchmarks. This workshop aims to provide a platform for both academic and industry researchers to engage in discussions, share insights, and foster collaboration toward achieving efficient and effective representation learning for multimodal information retrieval in the era of large foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14788
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
Fu, Junchen
Ge, Xuri
Xin, Xin
Yu, Haitao
Feng, Yue
Karatzoglou, Alexandros
Arapakis, Ioannis
Jose, Joemon M.
Information Retrieval
Multimodal representation learning has garnered significant attention in the AI community, largely due to the success of large pre-trained multimodal foundation models like LLaMA, GPT, Mistral, and CLIP. These models have achieved remarkable performance across various tasks of multimodal information retrieval (MIR), including web search, cross-modal retrieval, and recommender systems, etc. However, due to their enormous parameter sizes, significant efficiency challenges emerge across training, deployment, and inference stages when adapting these models' representation for IR tasks. These challenges present substantial obstacles to the practical adaptation of foundation models for representation learning in information retrieval tasks. To address these pressing issues, we propose organizing the first EReL@MIR workshop at the Web Conference 2025, inviting participants to explore novel solutions, emerging problems, challenges, efficiency evaluation metrics and benchmarks. This workshop aims to provide a platform for both academic and industry researchers to engage in discussions, share insights, and foster collaboration toward achieving efficient and effective representation learning for multimodal information retrieval in the era of large foundation models.
title The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2504.14788