Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Tianshi, Li, Fengling, Zhu, Lei, Li, Jingjing, Zhang, Zheng, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
von: Zhou, Hongyu, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyu, et al.
Veröffentlicht: (2025)
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
von: Zhan, Hao, et al.
Veröffentlicht: (2026)
von: Zhan, Hao, et al.
Veröffentlicht: (2026)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
von: Ma, Hongjian, et al.
Veröffentlicht: (2026)
von: Ma, Hongjian, et al.
Veröffentlicht: (2026)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
von: Yu, Mingzhe, et al.
Veröffentlicht: (2025)
von: Yu, Mingzhe, et al.
Veröffentlicht: (2025)
U-Sticker: A Large-Scale Multi-Domain User Sticker Dataset for Retrieval and Personalization
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
Learning Item Representations Directly from Multimodal Features for Effective Recommendation
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
von: Tran, Allie, et al.
Veröffentlicht: (2025)
von: Tran, Allie, et al.
Veröffentlicht: (2025)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
von: Sun, Huatuan, et al.
Veröffentlicht: (2025)
von: Sun, Huatuan, et al.
Veröffentlicht: (2025)
Small Stickers, Big Meanings: A Multilingual Sticker Semantic Understanding Dataset with a Gamified Approach
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
von: Chee, Heng Er Metilda, et al.
Veröffentlicht: (2025)
The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
Performance Evaluation in Multimedia Retrieval
von: Sauter, Loris, et al.
Veröffentlicht: (2024)
von: Sauter, Loris, et al.
Veröffentlicht: (2024)
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
von: Jia, Yanhao, et al.
Veröffentlicht: (2025)
von: Jia, Yanhao, et al.
Veröffentlicht: (2025)
VCR: Video representation for Contextual Retrieval
von: Nir, Oron, et al.
Veröffentlicht: (2024)
von: Nir, Oron, et al.
Veröffentlicht: (2024)
From ID-based to ID-free: Rethinking ID Effectiveness in Multimodal Collaborative Filtering Recommendation
von: Li, Guohao, et al.
Veröffentlicht: (2025)
von: Li, Guohao, et al.
Veröffentlicht: (2025)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
von: Nareti, Utsav Kumar, et al.
Veröffentlicht: (2024)
von: Nareti, Utsav Kumar, et al.
Veröffentlicht: (2024)
Multimodal Learned Sparse Retrieval for Image Suggestion
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
von: Ong, Rongqing Kenneth, et al.
Veröffentlicht: (2024)
von: Ong, Rongqing Kenneth, et al.
Veröffentlicht: (2024)
OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Agentic Mixed-Source Multi-Modal Misinformation Detection with Adaptive Test-Time Scaling
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
von: Tran, Quang-Linh, et al.
Veröffentlicht: (2025)
von: Tran, Quang-Linh, et al.
Veröffentlicht: (2025)
CM$^3$: Calibrating Multimodal Recommendation
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
von: Mo, Feng, et al.
Veröffentlicht: (2024)
von: Mo, Feng, et al.
Veröffentlicht: (2024)
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning
von: Zhang, Lingzi, et al.
Veröffentlicht: (2023)
von: Zhang, Lingzi, et al.
Veröffentlicht: (2023)
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
von: Gong, Ziyu, et al.
Veröffentlicht: (2025)
von: Gong, Ziyu, et al.
Veröffentlicht: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
von: Zhu, Shuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Shuqi, et al.
Veröffentlicht: (2024)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
von: Wu, Qiyu, et al.
Veröffentlicht: (2025)
von: Wu, Qiyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
von: Bin, Yi, et al.
Veröffentlicht: (2024) -
Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
von: Zhou, Hongyu, et al.
Veröffentlicht: (2025) -
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
von: Xu, Jinfeng, et al.
Veröffentlicht: (2025) -
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
von: Zhan, Hao, et al.
Veröffentlicht: (2026) -
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)