Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Kaiyan, Miao, Zhongtao, Tsuruoka, Yoshimasa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908475185954816
author Zhao, Kaiyan
Miao, Zhongtao
Tsuruoka, Yoshimasa
author_facet Zhao, Kaiyan
Miao, Zhongtao
Tsuruoka, Yoshimasa
contents Multimodal sentence embedding models typically leverage image-caption pairs in addition to textual data during training. However, such pairs often contain noise, including redundant or irrelevant information on either the image or caption side. To mitigate this issue, we propose MCSEO, a method that enhances multimodal sentence embeddings by incorporating fine-grained object-phrase alignment alongside traditional image-caption alignment. Specifically, MCSEO utilizes existing segmentation and object detection models to extract accurate object-phrase pairs, which are then used to optimize a contrastive learning objective tailored to object-phrase correspondence. Experimental results on semantic textual similarity (STS) tasks across different backbone models demonstrate that MCSEO consistently outperforms strong baselines, highlighting the significance of precise object-phrase alignment in multimodal representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
Zhao, Kaiyan
Miao, Zhongtao
Tsuruoka, Yoshimasa
Computation and Language
Multimodal sentence embedding models typically leverage image-caption pairs in addition to textual data during training. However, such pairs often contain noise, including redundant or irrelevant information on either the image or caption side. To mitigate this issue, we propose MCSEO, a method that enhances multimodal sentence embeddings by incorporating fine-grained object-phrase alignment alongside traditional image-caption alignment. Specifically, MCSEO utilizes existing segmentation and object detection models to extract accurate object-phrase pairs, which are then used to optimize a contrastive learning objective tailored to object-phrase correspondence. Experimental results on semantic textual similarity (STS) tasks across different backbone models demonstrate that MCSEO consistently outperforms strong baselines, highlighting the significance of precise object-phrase alignment in multimodal representation learning.
title Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
topic Computation and Language
url https://arxiv.org/abs/2508.00332