Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Hailang, Nie, Zhijie, Wang, Ziqiao, Shang, Ziyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal Retrieval
von: Liu, Yizhi, et al.
Veröffentlicht: (2025)
von: Liu, Yizhi, et al.
Veröffentlicht: (2025)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
Gradient-Guided Modality Decoupling for Missing-Modality Robustness
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
von: Huang, Shunyu, et al.
Veröffentlicht: (2026)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
von: Wang, Xue, et al.
Veröffentlicht: (2026)
von: Wang, Xue, et al.
Veröffentlicht: (2026)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
von: Kandhare, Mahesh, et al.
Veröffentlicht: (2024)
von: Kandhare, Mahesh, et al.
Veröffentlicht: (2024)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
Joint Explicit and Implicit Cross-Modal Interaction Network for Anterior Chamber Inflammation Diagnosis
von: Shao, Qian, et al.
Veröffentlicht: (2023)
von: Shao, Qian, et al.
Veröffentlicht: (2023)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
von: Sánchez, Jorge, et al.
Veröffentlicht: (2024)
von: Sánchez, Jorge, et al.
Veröffentlicht: (2024)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
Anisotropic Modality Align
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
von: Li, Yili, et al.
Veröffentlicht: (2025)
von: Li, Yili, et al.
Veröffentlicht: (2025)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
von: Luo, Bingjun, et al.
Veröffentlicht: (2025)
von: Luo, Bingjun, et al.
Veröffentlicht: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Provably Secure Robust Image Steganography via Cross-Modal Error Correction
von: Qi, Yuang, et al.
Veröffentlicht: (2024)
von: Qi, Yuang, et al.
Veröffentlicht: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025) -
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025) -
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
von: Zhang, Yafei, et al.
Veröffentlicht: (2025) -
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024) -
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)