Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Tianyu, Jung, Myong Chol, Clark, Jesse |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
von: Sogi, Naoya, et al.
Veröffentlicht: (2024)
von: Sogi, Naoya, et al.
Veröffentlicht: (2024)
CoopHash: Cooperative Learning of Multipurpose Descriptor and Contrastive Pair Generator via Variational MCMC Teaching for Supervised Image Hashing
von: Doan, Khoa D., et al.
Veröffentlicht: (2022)
von: Doan, Khoa D., et al.
Veröffentlicht: (2022)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
von: Fahim, Abrar, et al.
Veröffentlicht: (2024)
von: Fahim, Abrar, et al.
Veröffentlicht: (2024)
Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
Online Learning via Memory: Retrieval-Augmented Detector Adaptation
von: Jian, Yanan, et al.
Veröffentlicht: (2024)
von: Jian, Yanan, et al.
Veröffentlicht: (2024)
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
von: Liu, Zheyuan, et al.
Veröffentlicht: (2023)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2023)
Hierarchy-of-Visual-Words: a Learning-based Approach for Trademark Image Retrieval
von: Lourenço, Vítor N., et al.
Veröffentlicht: (2019)
von: Lourenço, Vítor N., et al.
Veröffentlicht: (2019)
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
Re-ranking the Context for Multimodal Retrieval Augmented Generation
von: Mortaheb, Matin, et al.
Veröffentlicht: (2025)
von: Mortaheb, Matin, et al.
Veröffentlicht: (2025)
PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models
von: Mishra-Sharma, Siddharth, et al.
Veröffentlicht: (2024)
von: Mishra-Sharma, Siddharth, et al.
Veröffentlicht: (2024)
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
von: Sarkar, Rohan, et al.
Veröffentlicht: (2024)
von: Sarkar, Rohan, et al.
Veröffentlicht: (2024)
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance
von: Mortaheb, Matin, et al.
Veröffentlicht: (2025)
von: Mortaheb, Matin, et al.
Veröffentlicht: (2025)
Sustainable transparency in Recommender Systems: Bayesian Ranking of Images for Explainability
von: Paz-Ruza, Jorge, et al.
Veröffentlicht: (2023)
von: Paz-Ruza, Jorge, et al.
Veröffentlicht: (2023)
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
Evidential Transformers for Improved Image Retrieval
von: Dordevic, Danilo, et al.
Veröffentlicht: (2024)
von: Dordevic, Danilo, et al.
Veröffentlicht: (2024)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
Embedding-based Retrieval in Multimodal Content Moderation
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
von: Levi, Hila, et al.
Veröffentlicht: (2024)
von: Levi, Hila, et al.
Veröffentlicht: (2024)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
AdaTask: A Task-aware Adaptive Learning Rate Approach to Multi-task Learning
von: Yang, Enneng, et al.
Veröffentlicht: (2022)
von: Yang, Enneng, et al.
Veröffentlicht: (2022)
Decoupled Training: Return of Frustratingly Easy Multi-Domain Learning
von: Wang, Ximei, et al.
Veröffentlicht: (2023)
von: Wang, Ximei, et al.
Veröffentlicht: (2023)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
von: Seo, Seonguk, et al.
Veröffentlicht: (2023)
von: Seo, Seonguk, et al.
Veröffentlicht: (2023)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
von: Xu, Yifan, et al.
Veröffentlicht: (2024)
von: Xu, Yifan, et al.
Veröffentlicht: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
von: Duan, Yue, et al.
Veröffentlicht: (2024)
von: Duan, Yue, et al.
Veröffentlicht: (2024)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
TelcoAI: Advancing 3GPP Technical Specification Search through Agentic Multi-Modal Retrieval-Augmented Generation
von: Ghosh, Rahul, et al.
Veröffentlicht: (2025)
von: Ghosh, Rahul, et al.
Veröffentlicht: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
von: Sarwar, Nobin
Veröffentlicht: (2025)
von: Sarwar, Nobin
Veröffentlicht: (2025)
Open Multimodal Retrieval-Augmented Factual Image Generation
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Multi-Label Plant Species Prediction with Metadata-Enhanced Multi-Head Vision Transformers
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2025)
von: Herasimchyk, Hanna, et al.
Veröffentlicht: (2025)
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
von: Xiao, Ling, et al.
Veröffentlicht: (2022)
von: Xiao, Ling, et al.
Veröffentlicht: (2022)
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
von: Gustineli, Murilo, et al.
Veröffentlicht: (2024)
von: Gustineli, Murilo, et al.
Veröffentlicht: (2024)
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark
von: Wasserman, Navve, et al.
Veröffentlicht: (2025)
von: Wasserman, Navve, et al.
Veröffentlicht: (2025)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
von: Du, Yang, et al.
Veröffentlicht: (2024)
von: Du, Yang, et al.
Veröffentlicht: (2024)
AutoPP: Towards Automated Product Poster Generation and Optimization
von: Fan, Jiahao, et al.
Veröffentlicht: (2025)
von: Fan, Jiahao, et al.
Veröffentlicht: (2025)
GBSK: Skeleton Clustering via Granular-ball Computing and Multi-Sampling for Large-Scale Data
von: Chen, Yewang, et al.
Veröffentlicht: (2025)
von: Chen, Yewang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
von: Sogi, Naoya, et al.
Veröffentlicht: (2024) -
CoopHash: Cooperative Learning of Multipurpose Descriptor and Contrastive Pair Generator via Variational MCMC Teaching for Supervised Image Hashing
von: Doan, Khoa D., et al.
Veröffentlicht: (2022) -
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
von: Alomari, Hani, et al.
Veröffentlicht: (2025) -
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023) -
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
von: Fahim, Abrar, et al.
Veröffentlicht: (2024)