Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915360172670976 |
|---|---|
| author | Alomari, Hani Sivakumar, Anushka Zhang, Andrew Thomas, Chris |
| author_facet | Alomari, Hani Sivakumar, Anushka Zhang, Andrew Thomas, Chris |
| contents | Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities. Set-based approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships. In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness. To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set. We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set. Our method achieves state-of-the-art performance on MS-COCO and Flickr30k without relying on external data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21538 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval Alomari, Hani Sivakumar, Anushka Zhang, Andrew Thomas, Chris Computer Vision and Pattern Recognition Information Retrieval Machine Learning Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities. Set-based approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships. In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness. To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set. We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set. Our method achieves state-of-the-art performance on MS-COCO and Flickr30k without relying on external data. |
| title | Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval |
| topic | Computer Vision and Pattern Recognition Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2506.21538 |