SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Siwei, Li, Yizhi, Zhu, Kang, Zhang, Ge, Liang, Yiming, Ma, Kaijing, Xiao, Chenghao, Zhang, Haoran, Yang, Bohao, Chen, Wenhu, Huang, Wenhao, Moubayed, Noura Al, Fu, Jie, Lin, Chenghua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911913134260224
author Wu, Siwei
Li, Yizhi
Zhu, Kang
Zhang, Ge
Liang, Yiming
Ma, Kaijing
Xiao, Chenghao
Zhang, Haoran
Yang, Bohao
Chen, Wenhu
Huang, Wenhao
Moubayed, Noura Al
Fu, Jie
Lin, Chenghua
author_facet Wu, Siwei
Li, Yizhi
Zhu, Kang
Zhang, Ge
Liang, Yiming
Ma, Kaijing
Xiao, Chenghao
Zhang, Haoran
Yang, Bohao
Chen, Wenhu
Huang, Wenhao
Moubayed, Noura Al
Fu, Jie
Lin, Chenghua
contents Multi-modal information retrieval (MMIR) is a rapidly evolving field, where significant progress, particularly in image-text pairing, has been made through advanced representation learning and cross-modality alignment research. However, current benchmarks for evaluating MMIR performance in image-text pairing within the scientific domain show a notable gap, where chart and table images described in scholarly language usually do not play a significant role. To bridge this gap, we develop a specialised scientific MMIR (SciMMIR) benchmark by leveraging open-access paper collections to extract data relevant to the scientific domain. This benchmark comprises 530K meticulously curated image-text pairs, extracted from figures and tables with detailed captions in scientific documents. We further annotate the image-text pairs with two-level subset-subcategory hierarchy annotations to facilitate a more comprehensive evaluation of the baselines. We conducted zero-shot and fine-tuning evaluations on prominent multi-modal image-captioning and visual language models, such as CLIP and BLIP. Our analysis offers critical insights for MMIR in the scientific domain, including the impact of pre-training and fine-tuning settings and the influence of the visual and textual encoders. All our data and checkpoints are publicly available at https://github.com/Wusiwei0410/SciMMIR.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13478
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
Wu, Siwei
Li, Yizhi
Zhu, Kang
Zhang, Ge
Liang, Yiming
Ma, Kaijing
Xiao, Chenghao
Zhang, Haoran
Yang, Bohao
Chen, Wenhu
Huang, Wenhao
Moubayed, Noura Al
Fu, Jie
Lin, Chenghua
Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Multi-modal information retrieval (MMIR) is a rapidly evolving field, where significant progress, particularly in image-text pairing, has been made through advanced representation learning and cross-modality alignment research. However, current benchmarks for evaluating MMIR performance in image-text pairing within the scientific domain show a notable gap, where chart and table images described in scholarly language usually do not play a significant role. To bridge this gap, we develop a specialised scientific MMIR (SciMMIR) benchmark by leveraging open-access paper collections to extract data relevant to the scientific domain. This benchmark comprises 530K meticulously curated image-text pairs, extracted from figures and tables with detailed captions in scientific documents. We further annotate the image-text pairs with two-level subset-subcategory hierarchy annotations to facilitate a more comprehensive evaluation of the baselines. We conducted zero-shot and fine-tuning evaluations on prominent multi-modal image-captioning and visual language models, such as CLIP and BLIP. Our analysis offers critical insights for MMIR in the scientific domain, including the impact of pre-training and fine-tuning settings and the influence of the visual and textual encoders. All our data and checkpoints are publicly available at https://github.com/Wusiwei0410/SciMMIR.
title SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
topic Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2401.13478