Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Alomari, Hani, Sivakumar, Anushka, Zhang, Andrew, Thomas, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
by: Sogi, Naoya, et al.
Published: (2024)
by: Sogi, Naoya, et al.
Published: (2024)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
by: Zhu, Tianyu, et al.
Published: (2024)
by: Zhu, Tianyu, et al.
Published: (2024)
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
by: Raja, Rahul, et al.
Published: (2025)
by: Raja, Rahul, et al.
Published: (2025)
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
by: Sarkar, Rohan, et al.
Published: (2024)
by: Sarkar, Rohan, et al.
Published: (2024)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)
by: Sivakumar, Anushka, et al.
Published: (2025)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Multi-event Video-Text Retrieval
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Automatic Creative Selection with Cross-Modal Matching
by: Kim, Alex, et al.
Published: (2024)
by: Kim, Alex, et al.
Published: (2024)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
by: Duan, Yue, et al.
Published: (2024)
by: Duan, Yue, et al.
Published: (2024)
Embedding-based Retrieval in Multimodal Content Moderation
by: Liang, Hanzhong, et al.
Published: (2025)
by: Liang, Hanzhong, et al.
Published: (2025)
Online Learning via Memory: Retrieval-Augmented Detector Adaptation
by: Jian, Yanan, et al.
Published: (2024)
by: Jian, Yanan, et al.
Published: (2024)
Evidential Transformers for Improved Image Retrieval
by: Dordevic, Danilo, et al.
Published: (2024)
by: Dordevic, Danilo, et al.
Published: (2024)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
by: Zhao, Pengcheng, et al.
Published: (2025)
by: Zhao, Pengcheng, et al.
Published: (2025)
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
by: Nie, Zhanheng, et al.
Published: (2025)
by: Nie, Zhanheng, et al.
Published: (2025)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
by: Du, Yang, et al.
Published: (2024)
by: Du, Yang, et al.
Published: (2024)
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
by: Levi, Hila, et al.
Published: (2024)
by: Levi, Hila, et al.
Published: (2024)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
by: Seo, Seonguk, et al.
Published: (2023)
by: Seo, Seonguk, et al.
Published: (2023)
Hierarchy-of-Visual-Words: a Learning-based Approach for Trademark Image Retrieval
by: Lourenço, Vítor N., et al.
Published: (2019)
by: Lourenço, Vítor N., et al.
Published: (2019)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
Image Fusion for Cross-Domain Sequential Recommendation
by: Wu, Wangyu, et al.
Published: (2024)
by: Wu, Wangyu, et al.
Published: (2024)
A Dataset and Framework for Learning State-invariant Object Representations
by: Sarkar, Rohan, et al.
Published: (2024)
by: Sarkar, Rohan, et al.
Published: (2024)
Region-Point Joint Representation for Effective Trajectory Similarity Learning
by: Long, Hao, et al.
Published: (2025)
by: Long, Hao, et al.
Published: (2025)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
by: Liu, Zheyuan, et al.
Published: (2023)
by: Liu, Zheyuan, et al.
Published: (2023)
Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
Re-ranking the Context for Multimodal Retrieval Augmented Generation
by: Mortaheb, Matin, et al.
Published: (2025)
by: Mortaheb, Matin, et al.
Published: (2025)
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance
by: Mortaheb, Matin, et al.
Published: (2025)
by: Mortaheb, Matin, et al.
Published: (2025)
Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
by: Moummad, Ilyass, et al.
Published: (2025)
by: Moummad, Ilyass, et al.
Published: (2025)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Rethinking Sparse Lexical Representations for Image Retrieval in the Age of Rising Multi-Modal Large Language Models
by: Nakata, Kengo, et al.
Published: (2024)
by: Nakata, Kengo, et al.
Published: (2024)
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
by: Pegia, Maria-Eirini, et al.
Published: (2026)
by: Pegia, Maria-Eirini, et al.
Published: (2026)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
Open Multimodal Retrieval-Augmented Factual Image Generation
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
by: Sun, Chengsong, et al.
Published: (2025)
by: Sun, Chengsong, et al.
Published: (2025)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
by: Sun, Zengbao, et al.
Published: (2024)
by: Sun, Zengbao, et al.
Published: (2024)
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
by: Fahim, Abrar, et al.
Published: (2024)
by: Fahim, Abrar, et al.
Published: (2024)
Similar Items
-
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
by: Sogi, Naoya, et al.
Published: (2024) -
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
by: Zhu, Tianyu, et al.
Published: (2024) -
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
by: Raja, Rahul, et al.
Published: (2025) -
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
by: Sarkar, Rohan, et al.
Published: (2024) -
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)