Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xiang, Li, Xuemei, Fang, Lexin, Zhang, Caiming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEMO: A Statistical Perspective for Efficient Image-Text Matching
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
by: Sun, Zengbao, et al.
Published: (2024)
by: Sun, Zengbao, et al.
Published: (2024)
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Automatic Creative Selection with Cross-Modal Matching
by: Kim, Alex, et al.
Published: (2024)
by: Kim, Alex, et al.
Published: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Rethinking Sparse Lexical Representations for Image Retrieval in the Age of Rising Multi-Modal Large Language Models
by: Nakata, Kengo, et al.
Published: (2024)
by: Nakata, Kengo, et al.
Published: (2024)
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval
by: Zhang, Zhuocheng, et al.
Published: (2026)
by: Zhang, Zhuocheng, et al.
Published: (2026)
Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning
by: Kühn, Paul Julius, et al.
Published: (2026)
by: Kühn, Paul Julius, et al.
Published: (2026)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
by: Inadumi, Shun, et al.
Published: (2025)
by: Inadumi, Shun, et al.
Published: (2025)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
by: Zhang, Weihang, et al.
Published: (2025)
by: Zhang, Weihang, et al.
Published: (2025)
Entity Image and Mixed-Modal Image Retrieval Datasets
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
by: Li, Haiwen, et al.
Published: (2024)
by: Li, Haiwen, et al.
Published: (2024)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
by: Mounis, Mohamed Darwish, et al.
Published: (2026)
by: Mounis, Mohamed Darwish, et al.
Published: (2026)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
by: Sun, Chengsong, et al.
Published: (2025)
by: Sun, Chengsong, et al.
Published: (2025)
Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation
by: Fang, Lexin, et al.
Published: (2025)
by: Fang, Lexin, et al.
Published: (2025)
Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
by: Wang, Junyi, et al.
Published: (2025)
by: Wang, Junyi, et al.
Published: (2025)
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
by: Sogi, Naoya, et al.
Published: (2024)
by: Sogi, Naoya, et al.
Published: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
by: Wang, Hongyi, et al.
Published: (2025)
by: Wang, Hongyi, et al.
Published: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
by: Alomari, Hani, et al.
Published: (2025)
by: Alomari, Hani, et al.
Published: (2025)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
by: Yang, Jiashu, et al.
Published: (2026)
by: Yang, Jiashu, et al.
Published: (2026)
A Novel Evaluation Framework for Image2Text Generation
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Modality-Balanced Learning for Multimedia Recommendation
by: Zhang, Jinghao, et al.
Published: (2024)
by: Zhang, Jinghao, et al.
Published: (2024)
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
by: Fahim, Abrar, et al.
Published: (2024)
by: Fahim, Abrar, et al.
Published: (2024)
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
Closing the Modality Gap for Mixed Modality Search
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
RED: Robust Event-Guided Motion Deblurring with Modality-Specific Disentanglement
by: Leng, Yihong, et al.
Published: (2025)
by: Leng, Yihong, et al.
Published: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
by: Zhou, Yinan, et al.
Published: (2025)
by: Zhou, Yinan, et al.
Published: (2025)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
by: Han, Haochen, et al.
Published: (2024)
by: Han, Haochen, et al.
Published: (2024)
Similar Items
-
DEMO: A Statistical Perspective for Efficient Image-Text Matching
by: Zhang, Fan, et al.
Published: (2024) -
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
by: Sun, Zengbao, et al.
Published: (2024) -
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
by: Zhang, Qi, et al.
Published: (2025) -
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024) -
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
by: Wang, Xin, et al.
Published: (2024)