DEMO: A Statistical Perspective for Efficient Image-Text Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Fan, Hua, Xian-Sheng, Chen, Chong, Luo, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
by: Zhang, Weihang, et al.
Published: (2025)
by: Zhang, Weihang, et al.
Published: (2025)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
by: Ma, Xiang, et al.
Published: (2024)
by: Ma, Xiang, et al.
Published: (2024)
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval
by: Zhang, Zhuocheng, et al.
Published: (2026)
by: Zhang, Zhuocheng, et al.
Published: (2026)
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
by: Byun, Jaeseok, et al.
Published: (2024)
by: Byun, Jaeseok, et al.
Published: (2024)
A Novel Evaluation Framework for Image2Text Generation
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization
by: Mao, Chen, et al.
Published: (2024)
by: Mao, Chen, et al.
Published: (2024)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
by: Yang, Jiashu, et al.
Published: (2026)
by: Yang, Jiashu, et al.
Published: (2026)
MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
POBEVM: Real-time Video Matting via Progressively Optimize the Target Body and Edge
by: Xian, Jianming
Published: (2024)
by: Xian, Jianming
Published: (2024)
Learning Positional Attention for Sequential Recommendation
by: Luo, Fan, et al.
Published: (2024)
by: Luo, Fan, et al.
Published: (2024)
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification
by: Mao, Chen, et al.
Published: (2024)
by: Mao, Chen, et al.
Published: (2024)
EndoFinder: Online Image Retrieval for Explainable Colorectal Polyp Diagnosis
by: Yang, Ruijie, et al.
Published: (2024)
by: Yang, Ruijie, et al.
Published: (2024)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
by: Khaertdinov, Bulat, et al.
Published: (2025)
by: Khaertdinov, Bulat, et al.
Published: (2025)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
by: Aiger, Dror, et al.
Published: (2025)
by: Aiger, Dror, et al.
Published: (2025)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
by: Sun, Zengbao, et al.
Published: (2024)
by: Sun, Zengbao, et al.
Published: (2024)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
by: Xu, Yexing, et al.
Published: (2026)
by: Xu, Yexing, et al.
Published: (2026)
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
Automatic Creative Selection with Cross-Modal Matching
by: Kim, Alex, et al.
Published: (2024)
by: Kim, Alex, et al.
Published: (2024)
Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
by: Choi, Wonseok, et al.
Published: (2025)
by: Choi, Wonseok, et al.
Published: (2025)
Iterative Optimal Attention and Local Model for Single Image Rain Streak Removal
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Chain-of-Thought Re-ranking for Image Retrieval Tasks
by: Wu, Shangrong, et al.
Published: (2025)
by: Wu, Shangrong, et al.
Published: (2025)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2024)
by: Xiao, Jian, et al.
Published: (2024)
Rethinking Detection Based Table Structure Recognition for Visually Rich Document Images
by: Xiao, Bin, et al.
Published: (2023)
by: Xiao, Bin, et al.
Published: (2023)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
by: Jiang, Haoyu, et al.
Published: (2024)
by: Jiang, Haoyu, et al.
Published: (2024)
Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
by: Wang, Junyi, et al.
Published: (2025)
by: Wang, Junyi, et al.
Published: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
Similar Items
-
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
by: Zhang, Qi, et al.
Published: (2025) -
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
by: Zhang, Weihang, et al.
Published: (2025) -
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
by: Ge, Xuri, et al.
Published: (2024) -
Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
by: Ma, Xiang, et al.
Published: (2024) -
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
by: Li, Da, et al.
Published: (2025)