A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Haochen, Luo, Minnan, Liu, Huan, Nan, Fang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
von: Duan, Yue, et al.
Veröffentlicht: (2024)
von: Duan, Yue, et al.
Veröffentlicht: (2024)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Interactive Multi-Turn Retrieval for Health Videos
von: Wu, Chengzheng, et al.
Veröffentlicht: (2026)
von: Wu, Chengzheng, et al.
Veröffentlicht: (2026)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
von: Fu, Junchen, et al.
Veröffentlicht: (2026)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
von: Li, Po-han, et al.
Veröffentlicht: (2024)
von: Li, Po-han, et al.
Veröffentlicht: (2024)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
von: Li, Minghan, et al.
Veröffentlicht: (2026)
von: Li, Minghan, et al.
Veröffentlicht: (2026)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
von: Liu, Han, et al.
Veröffentlicht: (2025)
von: Liu, Han, et al.
Veröffentlicht: (2025)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
von: Lian, Niu, et al.
Veröffentlicht: (2025)
von: Lian, Niu, et al.
Veröffentlicht: (2025)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
VKIE: The Application of Key Information Extraction on Video Text
von: An, Siyu, et al.
Veröffentlicht: (2023)
von: An, Siyu, et al.
Veröffentlicht: (2023)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
von: Shih, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shih, Yu-Fei, et al.
Veröffentlicht: (2025)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
A Comprehensive Survey on Composed Image Retrieval
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
A Survey of Multimodal Composite Editing and Retrieval
von: Li, Suyan, et al.
Veröffentlicht: (2024)
von: Li, Suyan, et al.
Veröffentlicht: (2024)
VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
von: Lokmanoglu, Ayse D, et al.
Veröffentlicht: (2025)
von: Lokmanoglu, Ayse D, et al.
Veröffentlicht: (2025)
TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations
von: Cakmak, Mert Can, et al.
Veröffentlicht: (2025)
von: Cakmak, Mert Can, et al.
Veröffentlicht: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Efficient Self-Supervised Video Hashing with Selective State Spaces
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022) -
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024) -
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
von: Luo, Tianci, et al.
Veröffentlicht: (2026) -
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025) -
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)