Compositional Image-Text Matching and Retrieval by Grounding Entities
Fuente:
arXiv
Saved in:
| Main Authors: | Vongala, Madhukar Reddy, Srivastava, Saurabh, Košecká, Jana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Spatial Reasoning with Open Vocabulary Object Detectors
by: Nejatishahidin, Negar, et al.
Published: (2024)
by: Nejatishahidin, Negar, et al.
Published: (2024)
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023)
by: Rajabi, Navid, et al.
Published: (2023)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024)
by: Rajabi, Navid, et al.
Published: (2024)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
by: Rajabi, Navid, et al.
Published: (2025)
by: Rajabi, Navid, et al.
Published: (2025)
PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting
by: Tran, Anh Thuan, et al.
Published: (2026)
by: Tran, Anh Thuan, et al.
Published: (2026)
VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM
by: Tran, Anh Thuan, et al.
Published: (2026)
by: Tran, Anh Thuan, et al.
Published: (2026)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Multi-temporal Adaptive Red-Green-Blue and Long-Wave Infrared Fusion for You Only Look Once-Based Landmine Detection from Unmanned Aerial Systems
by: Gallagher, James E., et al.
Published: (2025)
by: Gallagher, James E., et al.
Published: (2025)
GroundingBooth: Grounding Text-to-Image Customization
by: Xiong, Zhexiao, et al.
Published: (2024)
by: Xiong, Zhexiao, et al.
Published: (2024)
Entity Image and Mixed-Modal Image Retrieval Datasets
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
by: Blaga, Cristian-Ioan, et al.
Published: (2025)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
by: Bill, Eric Tillmann, et al.
Published: (2025)
by: Bill, Eric Tillmann, et al.
Published: (2025)
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
by: Ji, Zhong, et al.
Published: (2023)
by: Ji, Zhong, et al.
Published: (2023)
Grounding Image Matching in 3D with MASt3R
by: Leroy, Vincent, et al.
Published: (2024)
by: Leroy, Vincent, et al.
Published: (2024)
Text-based Aerial-Ground Person Retrieval
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Descriptive Image-Text Matching with Graded Contextual Similarity
by: Jang, Jinhyun, et al.
Published: (2025)
by: Jang, Jinhyun, et al.
Published: (2025)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
Anatomy-Aware Conditional Image-Text Retrieval
by: Zheng, Meng, et al.
Published: (2025)
by: Zheng, Meng, et al.
Published: (2025)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
by: Ouyang, Pengxiang, et al.
Published: (2025)
by: Ouyang, Pengxiang, et al.
Published: (2025)
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
Caption-Matching: A Multimodal Approach for Cross-Domain Image Retrieval
by: Iijima, Lucas, et al.
Published: (2024)
by: Iijima, Lucas, et al.
Published: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
by: Yang, Mingyue, et al.
Published: (2025)
by: Yang, Mingyue, et al.
Published: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024)
by: Miranda, Imanol, et al.
Published: (2024)
TPIE: Topology-Preserved Image Editing With Text Instructions
by: Jayakumar, Nivetha, et al.
Published: (2024)
by: Jayakumar, Nivetha, et al.
Published: (2024)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
Similar Items
-
Structured Spatial Reasoning with Open Vocabulary Object Detectors
by: Nejatishahidin, Negar, et al.
Published: (2024) -
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM
by: Rajabi, Navid, et al.
Published: (2024) -
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
by: Rajabi, Navid, et al.
Published: (2023) -
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
by: Rajabi, Navid, et al.
Published: (2024) -
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)