Caption-Matching: A Multimodal Approach for Cross-Domain Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Iijima, Lucas, Giakoumoglou, Nikolaos, Stathaki, Tania |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cluster Contrast for Unsupervised Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
SynCo: Synthetic Hard Negatives for Contrastive Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
Relational Representation Distillation
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
Discriminative and Consistent Representation Distillation
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
A Review on Discriminative Self-supervised Learning Methods in Computer Vision
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
Distilling Invariant Representations with Dual Augmentation
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
Unsupervised Training of Vision Transformers with Synthetic Negatives
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
Size Aware Cross-shape Scribble Supervision for Medical Image Segmentation
by: Yuan, Jing, et al.
Published: (2024)
by: Yuan, Jing, et al.
Published: (2024)
Towards Data-Efficient Medical Imaging: A Generative and Semi-Supervised Framework
by: Ma, Mosong, et al.
Published: (2025)
by: Ma, Mosong, et al.
Published: (2025)
Image edge enhancement for effective image classification
by: Bu, Tianhao, et al.
Published: (2024)
by: Bu, Tianhao, et al.
Published: (2024)
Enhanced Detection of Tiny Objects in Aerial Images
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
G-FARS: Gradient-Field-based Auto-Regressive Sampling for 3D Part Grouping
by: Cheng, Junfeng, et al.
Published: (2024)
by: Cheng, Junfeng, et al.
Published: (2024)
Comparing ImageNet Pre-training with Digital Pathology Foundation Models for Whole Slide Image-Based Survival Analysis
by: Papadopoulos, Kleanthis Marios, et al.
Published: (2024)
by: Papadopoulos, Kleanthis Marios, et al.
Published: (2024)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
Mean Height Aided Post-Processing for Pedestrian Detection
by: Yuan, Jing, et al.
Published: (2024)
by: Yuan, Jing, et al.
Published: (2024)
No Masks Needed: Explainable AI for Deriving Segmentation from Classification
by: Ma, Mosong, et al.
Published: (2025)
by: Ma, Mosong, et al.
Published: (2025)
CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization
by: Kritikos, Antonios, et al.
Published: (2026)
by: Kritikos, Antonios, et al.
Published: (2026)
Image compositing is all you need for data augmentation
by: Shermaine, Ang Jia Ning, et al.
Published: (2025)
by: Shermaine, Ang Jia Ning, et al.
Published: (2025)
DiffusionPrint: Learning Generative Fingerprints for Diffusion-Based Inpainting Localization
by: Giakoumoglou, Paschalis, et al.
Published: (2026)
by: Giakoumoglou, Paschalis, et al.
Published: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
SAGI: Semantically Aligned and Uncertainty Guided AI Image Inpainting
by: Giakoumoglou, Paschalis, et al.
Published: (2025)
by: Giakoumoglou, Paschalis, et al.
Published: (2025)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
by: Zhang, Junzhe, et al.
Published: (2024)
by: Zhang, Junzhe, et al.
Published: (2024)
Image Generation from Image Captioning -- Invertible Approach
by: Menon, Nandakishore S, et al.
Published: (2024)
by: Menon, Nandakishore S, et al.
Published: (2024)
ReMatch: Boosting Representation through Matching for Multimodal Retrieval
by: Liu, Qianying, et al.
Published: (2025)
by: Liu, Qianying, et al.
Published: (2025)
Large Language Models for Captioning and Retrieving Remote Sensing Images
by: Silva, João Daniel, et al.
Published: (2024)
by: Silva, João Daniel, et al.
Published: (2024)
RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
by: Gu, Jinjing, et al.
Published: (2025)
by: Gu, Jinjing, et al.
Published: (2025)
Temporal Image Caption Retrieval Competition -- Description and Results
by: Pokrywka, Jakub, et al.
Published: (2024)
by: Pokrywka, Jakub, et al.
Published: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Unsupervised Cross-Domain Image Retrieval via Prototypical Optimal Transport
by: Li, Bin, et al.
Published: (2024)
by: Li, Bin, et al.
Published: (2024)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
by: Mahmoud, Anas, et al.
Published: (2023)
by: Mahmoud, Anas, et al.
Published: (2023)
Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
by: Wang, Mengyu, et al.
Published: (2025)
by: Wang, Mengyu, et al.
Published: (2025)
Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
A Semantic Segmentation-guided Approach for Ground-to-Aerial Image Matching
by: Pro, Francesco, et al.
Published: (2024)
by: Pro, Francesco, et al.
Published: (2024)
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
by: You, Xiaoxing, et al.
Published: (2025)
by: You, Xiaoxing, et al.
Published: (2025)
RobustEMD: Domain Robust Matching for Cross-domain Few-shot Medical Image Segmentation
by: Zhu, Yazhou, et al.
Published: (2024)
by: Zhu, Yazhou, et al.
Published: (2024)
Similar Items
-
Cluster Contrast for Unsupervised Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2025) -
SynCo: Synthetic Hard Negatives for Contrastive Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2024) -
Relational Representation Distillation
by: Giakoumoglou, Nikolaos, et al.
Published: (2024) -
Discriminative and Consistent Representation Distillation
by: Giakoumoglou, Nikolaos, et al.
Published: (2024) -
A Review on Discriminative Self-supervised Learning Methods in Computer Vision
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)