MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Zhicheng, Shi, Qingyang, Lu, Jiasheng, Liang, Yingshan, Zhang, Xinyu, Wang, Yiran, Qin, Peiwu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912607736168448
author Du, Zhicheng
Shi, Qingyang
Lu, Jiasheng
Liang, Yingshan
Zhang, Xinyu
Wang, Yiran
Qin, Peiwu
author_facet Du, Zhicheng
Shi, Qingyang
Lu, Jiasheng
Liang, Yingshan
Zhang, Xinyu
Wang, Yiran
Qin, Peiwu
contents The multimodal relevance metric is usually borrowed from the embedding ability of pretrained contrastive learning models for bimodal data, which is used to evaluate the correlation between cross-modal data (e.g., CLIP). However, the commonly used evaluation metrics are only suitable for the associated analysis between two modalities, which greatly limits the evaluation of multimodal similarity. Herein, we propose MAJORScore, a brand-new evaluation metric for the relevance of multiple modalities ($N$ modalities, $N\ge3$) via multimodal joint representation for the first time. The ability of multimodal joint representation to integrate multiple modalities into the same latent space can accurately represent different modalities at one scale, providing support for fair relevance scoring. Extensive experiments have shown that MAJORScore increases by 26.03%-64.29% for consistent modality and decreases by 13.28%-20.54% for inconsistence compared to existing methods. MAJORScore serves as a more reliable metric for evaluating similarity on large-scale multimodal datasets and multimodal model performance evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
Du, Zhicheng
Shi, Qingyang
Lu, Jiasheng
Liang, Yingshan
Zhang, Xinyu
Wang, Yiran
Qin, Peiwu
Computer Vision and Pattern Recognition
Artificial Intelligence
The multimodal relevance metric is usually borrowed from the embedding ability of pretrained contrastive learning models for bimodal data, which is used to evaluate the correlation between cross-modal data (e.g., CLIP). However, the commonly used evaluation metrics are only suitable for the associated analysis between two modalities, which greatly limits the evaluation of multimodal similarity. Herein, we propose MAJORScore, a brand-new evaluation metric for the relevance of multiple modalities ($N$ modalities, $N\ge3$) via multimodal joint representation for the first time. The ability of multimodal joint representation to integrate multiple modalities into the same latent space can accurately represent different modalities at one scale, providing support for fair relevance scoring. Extensive experiments have shown that MAJORScore increases by 26.03%-64.29% for consistent modality and decreases by 13.28%-20.54% for inconsistence compared to existing methods. MAJORScore serves as a more reliable metric for evaluating similarity on large-scale multimodal datasets and multimodal model performance evaluation.
title MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.21365